LLMGPT2 - Tiny GPT-2: Train & Export to GGUF (C# / ILGPU / OpenCL / CUDA / CPU)

#145 · open · 0 comments

View on GitHub ↗

virex-84

You can add this fully open-source project. A C# implementation for creating and training extremely small GPT-2 style models from scratch, accelerated via OpenCL(Integrated AMD and Intel graphics)/CUDA(Nvidia) through ILGPU. The final exported GGUF model is only ~440 KB https://github.com/virex-84/LLMGPT2

Comments