Has Opus 5.5 Been Nerfed Yet? Tracking Frontier Drift
For months, developers have whispered about silent model downgrades. A new benchmark called Livenerf is finally putting those rumors to a rigorous test....

Every engineer working with modern language models carries a quiet, nagging suspicion. This we've all shipped a prompt that felt razor-sharp on Monday. Oddly enough, only to watch it hallucinate or coast on Friday. For months, the developer community has swapped dark theories about silent nerfs. Providers allegedly quantize weights behind our backs, swap out underlying routers, or throttle token budgets to manage setup costs. Until now, every single conversation about frontier drift boiled down to vibes versus vibes.
That is finally changing. A new project called Livenerf has stepped up to bring actual data to the argument. By building a deterministic, append-only evaluation loop targeting newly dropped releases like Claude Opus 5.5, this benchmark attempts to answer a simple, ruthless question: does a model quietly degrade after launch? It uses a locked use, frozen prompts, and raw public logs to track drift over time, relying on the UK AI Security Institute’s Inspect framework to keep things entirely transparent.

Making AI deterministic is notoriously difficult. This sampling factors vanish into black-box APIs, and you cannot simply switch off internal thinking traces. To bypass this, the benchmark isolates variance by focusing heavily on a pre-screened panel of notoriously tricky queries. If a model drops speed or tightens token output a lot over a thirty-day window, the telemetry catches it. [IMAGE]
What fascinates me most isn't just the pursuit of raw accuracy scores, but how behavioral changes manifest. It the project data point out a telling reality: lower working effort shows up much more clearly in token volume than in simple pass-fail metrics. When a system gets lazy, it doesn't always fail a calculus puzzle; it just stops showing its work. Instead, it gives you a (and this is key) rushed, truncated response of a deep reasoning path. Instead, it gives you a rushed, truncated response of a deep reasoning path.
We need more of this base across the entire ecosystem. It blindly, too many builders construct critical products — oddly. On top of shifting sand, trusting that black-box APIs will remain frozen in time. That matters. Independent tracking like this strips away the promo spin and forces accountability. If a provider alters how a model runs under the hood, we deserve to know, at once.






