KellerJordan/modded-nanogpt
NanoGPT (124M) in 90 seconds
NanoGPT (124M) in 90 seconds
Muon is an optimizer for hidden layers in neural networks
CIFAR-10 speedruns: 94% in 2.6 seconds and 96% in 27 seconds
The adaptation behavior of BatchNorm is no different than Norm-Free
EG plus/minus optimizer implemented in PyTorch
Code release for REPAIR: REnormalizing Permuted Activations for Interpolation Repair
PyTorch implementation of residual networks trained on CIFAR-10 dataset (2017)
welcome to the learning zone
An evaluation of the robust accuracy of the CrossMax Ensemble technique (Fort et al., 2024)
A fast, effective data attribution method for neural networks in PyTorch
neural networks don't minimize loss [caution: probably due to batchnorm]
Hacking sklearn's t-SNE implementation to animate embedding process
A replication of "Adversarial Examples Are Not Bugs, They Are Features" https://arxiv.org/abs/1905.02175
Variant of cifar10-airbench which removes several tricks. Ideal for research
Implementations of a few Algorithms and Data Structures in various languages
Fast and easy to use CIFAR-10 dataloader
Optimization algorithm which fits a ResNet to CIFAR-10 5x faster than SGD / Adam (with terrible generalization)
Replication of "Auto-encoder Based Data Clustering" Song et al
real oldschool notebook how they did it in the old days with none of that newfangled BS
Train to 94% on CIFAR-10 in 4.4 seconds on a single A100
Javascript math sandbox from before I knew about matlab
The simplest, fastest repository for training/finetuning medium-sized GPTs.
LLM training in simple, raw C/CUDA
Implementation of TriMap dimensionality reduction in PyTorch
Code for training and evaluating the robustness of models using pixelated data