MikeTrendsTrends right now

⬢github Python · 713 ★ +542 since we first saw it · pushed 20 h ago

ninjahawk/livenerf

Benchmark for tracking model capability after release.

livenerf is a 30-day, append-only benchmark testing whether Claude Opus 5.5 quietly degrades ('gets nerfed') after its 2026-09-22 launch. It runs daily via headless Claude Code on a frozen panel of 78 hard questions selected for inconsistent results, tracks drift statistically using Inspect AI and Anthropic's error-bars methodology, and logs everything raw. First results are expected around day 20.

Why now: It's getting attention because it addresses widely circulated community claims of post-launch model degradation with a pre-registered, day-0-baseline experiment instead of anecdote — and the daily series is actively running right now.

Who it is for: AI researchers, eval engineers, and anyone tracking whether frontier model quality changes after release.

llmbenchmarkevalspythonanthropicstatistics

Open on GitHub →

Stars over our 69 snapshots: 171 to 713, since 17 h ago.

Where people talked about it

API: https://socialmediatrends-api.osmike.com/v1/repos/ninjahawk/livenerf