Hey!
I'm currently researching BERT and I'm a bit confused about the position embeddings. I've come across many articles, websites, and blogs, but they all seem to say different things. Some claim that BERT uses learnable position embeddings, while others suggest it uses sin/cosine functions like the original Transformer model. There are also some sources that mention multiple ways to construct positional embeddings. Does anyone have a clear explanation on this? Also, if it’s the learnable positional embeddings, can anyone recommend some useful articles on the topic? I’ve had trouble finding any solid references myself
Thanks in advance!
In short, BERT uses learnable positional encodings.
This is a quote verbatim from the actual paper.
> Longer sequences are disproportionately expensive because attention is quadratic to the sequence length. To speed up pretraing in our experiments, we pre-train the model with sequence length of 128 for 90% of the steps. Then, we train the rest 10% of the **steps of sequence of 512 to learn the positional embeddings**.
Nevertheless, when you fine-tune BERT, the best practice would be to freeze positional encodings and don't update those variables.