Misleading "Uses tensor cores" print and dead TF32 setting in ch05/10_llm-training-speed

#1086 · closed · 2 comments

View on GitHub ↗

hellozdp

### Bug description Two small issues in `ch05/10_llm-training-speed/01_opt_single_gpu.py` (the second one also applies to `02_opt_multi_gpu_ddp.py`). Line numbers below are against the current `main` (`01_opt_single_gpu.py` blob sha `b99c969ad3`). ### 1. `Uses tensor cores:` reports CUDA availability, not tensor-core support `01_opt_single_gpu.py:380-389` ```python if torch.cuda.is_available(): capability = torch.cuda.get_device_capability() if capability[0] >= 7: torch.set_float32_matmul_precision("high") print("Uses tensor cores") else: print("Tensor cores not supported on this GPU. Using default precision.") print(f"Uses tensor cores: {torch.cuda.is_available()}") # line 389 ``` On a pre-Volta card (e.g. GTX 1080, compute capability 6.1) the output contradicts itself: ``` Tensor cores not supported on this GPU. Using default precision. Uses tensor cores: True ``` Suggested fix: ```python has_tensor_cores = ( torch.cuda.is_available() and torch.cuda.get_device_capability()[0] >= 7 ) print(f"Uses tensor cores: {has_tensor_cores}") ``` `02_opt_multi_gpu_ddp.py` does not have this print, so this part is specific to `01`. ### 2. `set_float32_matmul_precision("high")` has no effect once the model is cast to bf16 Both scripts enable TF32 and then convert the whole model to bfloat16: | file | TF32 | bf16 cast | |---|---|---| | `01_opt_single_gpu.py` | line 385 | line 415 `model.to(device).to(torch.bfloat16)` | | `02_opt_multi_gpu_ddp.py` | line 455 | line 489 `model = model.to(torch.bfloat16)` | `set_float32_matmul_precision` only affects **fp32** matmuls, and after the cast there are none left. I checked this by hooking every `nn.Linear` in `GPTModel` after `.to(torch.bfloat16)`: ``` parameters: 37x torch.bfloat16 buffers: 2x torch.bfloat16 forward pass, 13 Linear layers: input bfloat16, weight bfloat16, output bfloat16 final logits: torch.bfloat16 ``` So as the scripts are currently written, the TF32 branch never does anything. It is still useful if a reader comments out the bf16 cast to compare "fp32 + TF32" against "pure bf16", but as written it reads as if the two optimizations stack, which they don't. A short comment would make the intent clear, e.g.: ```python # Only relevant if the bfloat16 cast below is disabled; with a bf16 model # there are no fp32 matmuls left for this setting to affect. torch.set_float32_matmul_precision("high") ``` Happy to open a PR for either or both if that would be useful. Thanks for the book and the repo — they have been a great resource. ### What operating system are you using? macOS ### Where do you run your code? Local (laptop, desktop) ### Environment ``` { "machine": "MacBook Pro (Mac15,6, Apple M3 Pro)", "OS": "macOS-15.7.3-arm64-arm-64bit", "architecture": "arm64", "Python": "3.11.15", "PyTorch": "2.13.0", "installed_with": "pip", "installation_source": null, "CUDA_build": null, "CUDA_available": false, "MPS_available": true } ``` Note: this is a code-reading report, not a runtime failure — both issues are visible from the source and neither requires a CUDA device to observe.

Comments

mira687

Both check out by reading, and #1 has a second half that survives your fix: the threshold itself is wrong. `set_float32_matmul_precision("high")` buys TF32, and TF32 starts at Ampere. PyTorch's doc for `torch.backends.cuda.matmul.allow_tf32`: "whether TensorFloat-32 tensor cores may be used in matrix multiplications on Ampere or newer GPUs"; and per `set_float32_matmul_precision`'s own docstring, with no fast algorithm available the matmul is computed "as if the precision is 'highest'". The gate is `capability[0] >= 7` and the comment beside it names Volta and Turing — both have tensor cores, neither has TF32. So on a V100 or a T4 `has_tensor_cores` still prints `True` while line 385 changes nothing. `>= 8` is the condition that matches what the call does. The bf16 cast sits on the same boundary, which makes the else branch misleading independently of the print: `01:415` and `02:489` cast unconditionally, so "Using default precision" never describes the run. PyTorch's own native-bf16 check is also `major >= 8` — `torch.cuda.is_bf16_supported()` returns True immediately at compute capability 8.0 and below that only via its `_check_bf16_tensor_supported` emulation fallback. So the 1080 in your example prints "Tensor cores not supported on this GPU. Using default precision." and then trains the model in bf16 anyway.

rasbt

Thanks so much. Very good catch. Not sure what I was thinking then. Fixing it via #1095