Hey,
In the paper it is stated that batch size is 350seconds, is that per GPU, or total batch size? Also, if I follow the hyper-parameters of the HuBERT repo to pre-train the model I get a batch size of 2800seconds on 32 V100 GPUs
Kind Regards,
Goksenin Yuksel
Same question here!
@root-goksenin BTW, in the WavLM paper, the authors said that the Base and Base+ models are fine-tuned on 8 GPUs with a batch size equivalent to 200 seconds of audio for each GPU. The Large model is fine-tuned on 24 GPUs with a batch size equivalent to 80 seconds of audio for each GPU. So, it might be **per GPU batch size** I guess?
@mindmapper15 I am also thinking that it may be per-gpu. This paper: https://arxiv.org/abs/2305.10005 shows that WavLM's batch size is 187 minutes, which matches 350 seconds * 32 GPUs