⬢github C++ · 35 ★ +34 since we first saw it · pushed 12 h ago · MIT
lostmsu/TurboGPT
Train a tiny GPT in under a minute (CUDA only)
TurboGPT is a tiny CUDA C++ implementation for training a byte-level GPT model extremely fast — a ~22KiB transformer trains in about 13 seconds. It checkpoints model, optimizer, scheduler, and trainer state, emits TensorBoard-compatible logs, and includes a verification test script.
Why now: It was shared on Hacker News as a Show HN demo, drawing attention for training a tiny transformer in seconds on a single GPU.
Who it is for: ML hobbyists and developers with NVIDIA GPUs who want to experiment with transformer training end-to-end without large compute budgets.
Stars over our 55 snapshots: 1 to 35, since 13 h ago.
Where people talked about it
API: https://socialmediatrends-api.osmike.com/v1/repos/lostmsu/TurboGPT