Inconsistency in LoRALayer scaling factor

#1079 · closed · 8 comments

View on GitHub ↗

frx-wintermute

### Bug description In appendix-E the LoRALayer class was [modified](https://github.com/rasbt/LLMs-from-scratch/blob/main/appendix-E/01_main-chapter-code/appendix-E.ipynb?short_path=9d798ed#L672) with respect to the original code found in the book. The scaling factor is no longer `self.alpha`, but `self.alpha/self.rank`: ``` x = (self.alpha / self.rank) * (x @ self.A @ self.B) ``` However, the [solution](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/01_main-chapter-code/exercise_experiments.py#L133) for exercise 7.4 still uses the original implementation of the LoRALayer class: ``` x = self.alpha * (x @ self.A @ self.B) ``` This creates a difference in the results. Since the code I was using followed the new implementation (which divides `self.alpha` by `self.rank`), I had to pass `alpha=256` to `replace_linear_with_lora()` in order to get finetuning results similar to those shown in the [output](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/01_main-chapter-code/exercise-solutions.ipynb?short_path=8dfe053#L969) of the solution... I think this inconsistency should be fixed in the Github repository. Please note that, with the scaling factor `(self.alpha / self.rank)`, setting `alpha=16` resulted in poor finetuning with significantly higher training and validation loss values. Setting `alpha=256` resulted in loss values very similar to those shown in the solution: ``` Ep 1 (Step 000000): Train loss 2.529, Val loss 2.537 [...] Ep 2 (Step 000230): Train loss 0.308, Val loss 0.654 ``` However, at the end of epoch 2, the sentence 'The chef cooks the meal every day.' is still not very well converted into passive form: ``` Below is an instruction that describes a task. Write a response that appropriately completes the request. ### Instruction: Convert the active sentence to passive: 'The chef cooks the meal every day.' ### Response: The meal was cooked by the chef.<|endoftext|>The following sentence should be in the active voice. ### Input: The chef cooked the meal every day. ### Response: The meal was cooked by ``` For comparison, the finetuning (from chapter 7) without LoRA gave me a much more accurate response: ``` Below is an instruction that describes a task. Write a response that appropriately completes the request. ### Instruction: Convert the active sentence to passive: 'The chef cooks the meal every day.' ### Response: The meal is cooked every day by the chef.<|endoftext|>The following is an instruction that describes a task. Write a response that appropriately completes the request. ### Instruction: What is the capital of the United Kingdom ``` Is there anything wrong with the code I am running? What could it be? I compared with the exercise solution, and it seems to me that the code I am running is equivalent... ### What operating system are you using? Linux ### Where do you run your code? Local (laptop, desktop) ### Environment ``` ```

Comments

mira687

Your `alpha=256` isn't an approximation, it's exact, and that answers your question. `ch07/01_main-chapter-code/exercise_experiments.py` never stores `self.rank` at all, so `replace_linear_with_lora(model, rank=16, alpha=16)` on line 429 gives an effective adapter scale of **16**. appendix-E's `(self.alpha / self.rank)` with those same two arguments gives **1**. Identical call, 16x apart. 256/16 = 16, so you reproduced the solution's scale precisely, and there's nothing wrong with the code you're running. Which means the scaling isn't what separates your two runs anymore. Worth a second look at the output too: your LoRA run does produce a correct passive, "The meal was cooked by the chef" - it shifts tense and drops "every day" against the full-finetune's "is cooked every day by the chef". The run-on after `<|endoftext|>` shows up in both of your pastes, so that part isn't LoRA either.

rasbt

Thanks for reporting, I am updating it here in https://github.com/rasbt/LLMs-from-scratch/pull/1080 to make it consistent with the updated appendix. As for the results, I reran it in both cases and got the same results: > The meal is cooked every day by the chef. So the issue you are observing could be due to rounding or other numerical differences like in #977 perhaps?

frx-wintermute

> Thanks for reporting, I am updating it here in [#1080](https://github.com/rasbt/LLMs-from-scratch/pull/1080) to make it consistent with the updated appendix. You're welcome, thanks to you for fixing the inconsistency on the Github repository! > > As for the results, I reran it in both cases and got the same results: > > > The meal is cooked every day by the chef. > > So the issue you are observing could be due to rounding or other numerical differences like in [#977](https://github.com/rasbt/LLMs-from-scratch/issues/977) perhaps? Well, I am running on CPU: do you mean that the numerical differences among various platforms (CPU, NVIDIA GPU, MPS), although small in terms of loss values, can lead to such a significant change in the output of the 'The chef cooks the meal every day.' prompt?

mira687

It can, and what looks like a big change starts as a single token. Both responses start `The meal`, then yours picks ` was` and the rerun picks ` is`. `generate_text_simple` is greedy, so if those two were close, small numerical differences in the trained weights are enough to flip that argmax, and the rest of the sentence is the model continuing from a different prefix. Matching losses don't rule it out: loss is averaged over every token, and this is one position. You can check it on your trained model by replaying the decode and printing the runner-up at each step. If ` is` is a close second at the step where ` was` wins, that's your answer. ```python model.eval() ids = text_to_token_ids(format_input(val_data[0]), tokenizer).to(device) with torch.no_grad(): for _ in range(10): probs = torch.softmax(model(ids)[0, -1], dim=-1) top = torch.topk(probs, 2) first, second = top.indices.tolist() print(repr(tokenizer.decode([first])), round(top.values[0].item(), 3), "| runner-up", repr(tokenizer.decode([second])), round(top.values[1].item(), 3)) ids = torch.cat([ids, top.indices[:1].view(1, 1)], dim=1) ```

frx-wintermute

> It can, and what looks like a big change starts as a single token. Both responses start `The meal`, then yours picks ` was` and the rerun picks ` is`. `generate_text_simple` is greedy, so if those two were close, small numerical differences in the trained weights are enough to flip that argmax, and the rest of the sentence is the model continuing from a different prefix. Matching losses don't rule it out: loss is averaged over every token, and this is one position. That makes sense, thanks for your insight. > > You can check it on your trained model by replaying the decode and printing the runner-up at each step. If ` is` is a close second at the step where ` was` wins, that's your answer. Well, it depends on your definition of "close": ``` '\n' 1.0 | runner-up '\n\n' 0.0 '\n' 1.0 | runner-up 'The' 0.0 '###' 1.0 | runner-up '#' 0.0 ' Response' 0.956 | runner-up ' Input' 0.043 ':' 1.0 | runner-up ' :' 0.0 '\n' 0.999 | runner-up '<|endoftext|>' 0.0 'The' 0.967 | runner-up 'He' 0.013 ' meal' 0.605 | runner-up ' chef' 0.322 ' was' 0.524 | runner-up ' is' 0.418 ' cooked' 0.591 | runner-up ' prepared' 0.387 ``` 0.418 not far away from 0.524, but can small numerical differences in the trained weights be enough to flip their ordering?

mira687

For weights that differ by rounding, no. For a separate training run, yes, and those aren't the same thing. Which of the two wins depends only on the gap between their logits, ln(0.524 / 0.418) ≈ 0.23, and running the same weights on another device changes logits by far less than that. Training is different: each step starts from the previous step's weights, so a tiny arithmetic difference early on feeds every gradient after it, and by step 230 your run and the rerun can be two different fits with nearly the same loss. At 52/42 the model hasn't really picked a tense, so that's the first place two fits would disagree. To measure it rather than take my word for it, retrain with a different `torch.manual_seed` and print the same table. A new seed is a much bigger nudge than CPU vs GPU, so if ` was` still wins by about the same margin, I'm wrong and it's something else.

frx-wintermute

> For weights that differ by rounding, no. For a separate training run, yes, and those aren't the same thing. [...] > To measure it rather than take my word for it, retrain with a different `torch.manual_seed` and print the same table. A new seed is a much bigger nudge than CPU vs GPU, so if ` was` still wins by about the same margin, I'm wrong and it's something else. Amazing! I started from the pre-trained model again and tried to finetune with a different seed (621 in stead of 123, for the record). It resulted in similar loss values: ``` Ep 1 (Step 000000): Train loss 2.575, Val loss 2.567 [...] Ep 2 (Step 000230): Train loss 0.327, Val loss 0.665 ``` But, at the end of epoch 2, the sentence 'The chef cooks the meal every day.' is accurately converted into passive form: ``` Below is an instruction that describes a task. Write a response that appropriately completes the request. ### Instruction: Convert the active sentence to passive: 'The chef cooks the meal every day.' ### Response: The meal is cooked every day by the chef.<|endoftext|>The following is an instruction that describes a task. Write a response that appropriately completes the request. ### Instruction: What is the boiling point of water in ``` The table shows that now the appropriate tokens are chosen, and, furthermore, they win by a significant margin: ``` '\n' 0.999 | runner-up '<|endoftext|>' 0.0 '\n' 1.0 | runner-up 'The' 0.0 '###' 1.0 | runner-up '##' 0.0 ' Response' 0.95 | runner-up ' Input' 0.049 ':' 1.0 | runner-up ' from' 0.0 '\n' 0.999 | runner-up " '" 0.001 'The' 0.984 | runner-up 'He' 0.003 ' meal' 0.669 | runner-up ' chef' 0.285 ' is' 0.724 | runner-up ' was' 0.11 ' cooked' 0.83 | runner-up ' prepared' 0.11 ``` Hence, the conclusion seems to be that small numerical differences (due to different computing platforms) or different (pseudo-)random numbers during the finetuning can lead to very different finetuned models. Maybe this is exacerbated by the use of a relatively small dataset for the finetuning... Thank you so much for the time you spent in investigating these results! I think it was useful to me, as it helped me to better understand what was going on. Bye and thanks again for your helpfulness!

mira687

Worth keeping the two halves apart, because the seed run only tested one of them. Seeds, yes. Devices, still no — nothing here measured that, and a 0.23 logit gap is orders of magnitude above device-level rounding. What strikes me is how little actually changed. Nine of your ten positions pick the same token. The only numbers that really moved are the three where run 1 was under 0.7: ` meal` 0.605→0.669, the tense flip, and ` cooked` 0.591→0.83. Everything already above 0.95 barely budged. That's not two very different models, it's one undetermined position resolving the other way. And it's undetermined for a findable reason. Your prompt is `instruction-data.json` index 1045 — the first item of `val_data`, since `gpt_instruction_finetuning.py:173-175` splits 935/110/55 sequentially — so it's held out, and the gold output is `The meal is cooked by the chef every day.` The training split contains exactly five present-tense→passive examples (`Julia throws`, `company employs`, `gardener waters`, `teacher explains`, `We celebrate`), all mapping to is/are, no counterexamples; every past-tense source maps to was/were. So the tense-copying rule is consistent in the data, just supported by five examples out of 935. Thin enough that whether a given run picks it up isn't settled.