Train a small LLM from scratch in two hours
Code and guides to pretrain and fine-tune a 64M-parameter language model on a single consumer GPU.
From the repository: “🧠 Train a 64M-parameter LLM from scratch in just 2h!”