Yhn SportFootball first seen 1 d ago, last 16 min ago, peak #1
Qwen 125B model runs fast on a single RTX 4090
Original: Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
A new open-source project called Strata claims to run Qwen 3.8 Flash Next, a 125-billion-parameter model, on a single consumer RTX 4090 graphics card at roughly 100 tokens per second. The claim has drawn strong interest from developers, who are discussing performance figures, memory requirements and whether the results hold up in practice.
Why now: Running a very large model at high speed on consumer hardware would make local AI far cheaper and more accessible.
Rank over time, top of the chart is #1. 172 snapshots from 1 d ago to 16 min ago.
Evidence
API: https://socialmediatrends-api.osmike.com/v1/trends/988765