Operators are again reporting that Claude “got dumber,” but the strongest read is not simple model regression. The issue appears to be production scaffolding, including cache TTL behavior, adaptive thinking, and effort-level flips, according to this Reddit analysis.
The enterprise risk is that model quality can shift without a model version change. CTOs should treat orchestration settings, cache policy, context handling, and inference effort as part of the production AI surface, not as plumbing.
The cache TTL angle is especially important because it links quality complaints to cost volatility. This mirrors Niels’ earlier Anthropic bill spike investigation: small infrastructure changes can affect both answer quality and unit economics.
The practical takeaway is to monitor prompts, outputs, latency, token use, cache hit rates, and reasoning-effort settings together. AI leaders should require regression tests for the full application stack, not just benchmark the base model.
Today’s theme: enterprise AI reliability is increasingly a systems engineering problem, not a model leaderboard problem.