⬢github C · 23.6K ★ +298 since we first saw it · pushed 15 d ago · MIT
antirez/ds4
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
DwarfStar (antirez/ds4) is a small, self-contained C inference engine for running a handful of specific open-weight LLMs—DeepSeek V4 Flash/PRO, GLM 5.x, Qwen3.8—on consumer hardware like 96GB+ Macs, DGX Spark, and Strix Halo. It supports Metal, CUDA (including multi-GPU setups like 8x L40S), and ROCm, with SSD streaming for machines short on RAM, tensor/pipeline parallelism, an HTTP server, and its own GGUF files.
Why now: It's gaining attention as antirez's specialized alternative to llama.cpp for a few curated models, with an open 'AI full disclosure' about heavy AI-agent-assisted development sparking discussion.
Who it is for: Hobbyists and developers with high-RAM consumer machines who want fast local inference of specific frontier open-weight models.
Stars over our 117 snapshots: 23.3K to 23.6K, since 1 d ago.
Where people talked about it
- ⬢github antirez/ds4 7 h ago
API: https://socialmediatrends-api.osmike.com/v1/repos/antirez/ds4