Autonomy Per Interruption
On August 5, 2025, a routing bug began sending some Claude Sonnet 4 requests to servers configured for the wrong context window. Two further infrastructure bugs shipped on August 25 and 26. For the following month the three degraded Claude’s output without failing it. Requests succeeded and answers came back worse. At the worst point, in late August, roughly 30% of Claude Code users had at least one message served by the wrong hardware, and the misrouting was sticky, so a session that landed on a bad server tended to stay there. There was no downtime, no incident banner, and no error for a retry to catch. Every uptime dashboard reported green throughout, correctly: what broke was not availability [1].
The pattern recurred. Between March and April 2026, three further issues — a changed reasoning default, a caching bug that dropped prior thinking on every turn, and a verbosity constraint added to the system prompt — degraded output on Sonnet 4.6, Opus 4.6, and Opus 4.7 across roughly six weeks. The postmortem records that the API was not impacted [2].
The incident cost interactive users little. Degradation was visible within a day, the remedy was to re-run the request, and the loss was measured in minutes. For scheduled work — a long task handed to a model and collected later — the same degradation went undetected for weeks and surfaced only in the output.
That asymmetry is the subject here. Benchmarks price intelligence per dollar and per hour, which is close to the right question for supervised work: a human is already in the loop, per-turn quality binds, and an interruption costs seconds. For unattended work the binding term is different — how long a model runs before a human must intervene, and how stable that figure is week to week. Call it autonomy per interruption; it is largely unmeasured.
Two distinct failures are being grouped under that heading, and they diverge as runs lengthen. A run can stop — the quota is exhausted, the endpoint returns a 529, the scheduler abandons it — or it can complete incorrectly, as the August runs did: full completion, plausible output, quietly degraded. The first is detected immediately and costs one run. The second contaminates everything built downstream before anyone notices.
It registers on no metric proposed here, including the one in the title. A silently degraded run has an interruption rate of zero.
Most of this friction originates not in the weights but in the platform around them: quotas, rate limits, capacity churn, and the harness. They are not separable at purchase: they ship as a bundle, and that’s what gets scheduled. Until interruption rate and quota variance are reported alongside capability, “best model” answers a question few unattended workloads are asking.
The Variance Term#
Interruption rate is only half the measure; the other half is variance. Scheduling breaks not on limits so much as on limits that move. A fixed quota is an engineering constraint. Capacity is provisioned around it and the pipeline runs. A quota that shifts forces provisioning against the worst observed week, so variance imposes a cost even in the weeks when the limit is generous.
Variance is also what the current model scorecards cannot see. In July 2025, Claude Code limits tightened sharply with no announcement; asked what had changed, Anthropic said only that it was aware some users were seeing slower responses, and its pricing documentation described tiers without guaranteeing any particular level of access [3]. That December, developers on Google’s own forum reported 429 errors on Gemini while inside their paid-tier limits — a thread running six months across seventeen participants, answered by Google staff with the observation that limits also apply per-minute and per-region, in ways the published table does not show [4]. In both cases the ceiling was not where the documentation placed it, and a schedule built on the documented figure would break. The August degradation was the other failure. The ceiling held, requests succeeded, and the answers were worse. The two share one property, which is that no dashboard reported either.
Where the Cost Lands#
The friction cost does not disappear; it moves to payroll. Some of the resulting scaffolding is ordinary distributed-systems hygiene — retries and checkpoints are table stakes against any external dependency. The marginal cost sits above that baseline — quota-aware schedulers, context-window management, fallback hierarchies, a standing rotation to monitor unattended runs — and it is an ongoing operation, not a one-time build. The API invoice remains unremarkable; the expense is engineer-hours spent making an unpredictable system schedulable. The friction term is therefore unmeasured at the provider and hidden at the buyer, absorbed into headcount where it is not attributed to the platform choice.
The scaffolding payroll buys then compounds the problem. Generic layers transfer — a retry wrapper works against any API — but the costliest components do not: schedulers tuned to one vendor’s quota rhythms, fallback chains tuned to its degradation behavior, prompt-level workarounds for particular failure modes. The more friction a platform imposes, the more of this accumulates, and the higher the switching cost. The asymmetry runs one way. The workarounds are sunk cost already capitalized, while leaving requires an alternative demonstrably better on a term nobody measures.
Friction operates as a retention mechanism.
The signal that should discipline the vendor therefore arrives attenuated. Were friction reflected in the API invoice, procurement would negotiate it; because it falls on payroll, it reaches the vendor as support tickets and anecdote rather than as price. A partial market does exist: provisioned-throughput products and priority tiers are insurance sold against the vendor’s own overload errors [5], [6]. Their existence indicates vendors regard friction as worth money, but they price throughput rather than interruptions or variance, they are gated to enterprise scale, and when demand rose, Anthropic’s priority tier stopped accepting new customers [6]. Below that tier the feedback loop that would price predictability is severed at the point where the cost is incurred. Laundering is not the only mechanism keeping friction unmeasured — the metric is difficult to standardize, and every party in the chain has somewhere else to point — but it is the load-bearing one.
Within the buyer, the last opportunity to surface the cost is lost to internal incentives. Maintenance of this kind is not what promotion rewards, so it is chronically understaffed, and when a run fails, responsibility attaches to the team or the model rather than the platform. Engineers can identify the cost precisely — it is in the repository, in the scheduler named after the quota — but it has no line item, so those who set budgets do not see it. The people with the clearest view of the cost are the people with no pricing power.
What Procurement Buys#
Because the cost is hidden in payroll, procurement optimizes the wrong term. Buyers select on the measured number, which is not wrong so much as partial. Placing reliability beside capability separates the two rankings. Across ten models and 23,392 episodes, capability and reliability diverge substantially, with multi-rank inversions appearing as task horizons lengthen [7]. The divergence disadvantages whoever buys the leader. In the same study, frontier models recorded the highest meltdown rates, up to 19%, because they attempt ambitious multi-step strategies that sometimes spiral [7]. The misallocation is not that spending follows the most capable model; it is that the spending is priced as though intelligence per dollar were the whole product.
The leading bid also carries the leader’s queue. Waitlists, quota reductions, and overload errors concentrate on whatever model has most recently been reviewed, just as contracts are signed. Speaking at Code with Claude, Dario Amodei said Anthropic had planned for 10x growth and observed closer to 80x in revenue and usage [8]. The response was two weekly caps, one general and one specific to the flagship: on the $100 plan, 15 to 35 hours of Opus 4 per week against 140 to 280 hours of Sonnet 4 [9]. Documentation advises switching to a lighter model on reaching a limit [10]. DeepSeek suspended API credit top-ups within weeks of topping the charts, existing balances remaining spendable [11]. In each case what is rationed is access, not quality; Anthropic states that it does not reduce model quality for demand, time of day, or server load [1].
Whether this generalizes is unclear: the rationing above is visible in channels bought directly and absent from the one channel with systematic public data. Ranked by measured availability on a reseller that has already purchased its capacity, expensive models are the steadier ones. Across the 293 models with a full history, the top price quintile spends roughly a tenth as much time below the registry’s health threshold as the bottom [12]. This cuts against the account above, and it is measured on the wrong side of the meter, since absorbing the queue is a reseller’s function. The defensible claim is narrower. Scarcity concentrates on the frontier where capacity is bought directly, and elsewhere it is not instrumented.
A reasonable objection is that this describes a moment rather than a structure — capacity arrives and the queue clears. It does clear, for that model. But friction is transient per model and persistent per position. Last year’s frontier is abundant and inexpensive; the current frontier is the rationed one. The queue relocates rather than dissolves, and it stands in front of whatever is most worth scheduling.
What Already Gets Measured#
The field is measuring part of this already: the model half. METR’s time-horizon work measures close to the right quantity. The task length a model completes unattended at a given success rate [13], [14]. The literature also has the appropriate consistency metric. τ-bench introduced passk, which asks whether an agent succeeds on all k attempts rather than on any of them, and the decline is steep: the GPT-4o configuration that solved 61% of τ-retail tasks on a single attempt solved roughly 25% of them eight times consecutively [15]. Vendor system cards report pass@k instead — the probability of succeeding at least once in k attempts. The Claude Sonnet 4.6 card reports pass@k on its agentic benchmarks and no passk anywhere [16]. Of the two available metrics, the industry publishes the one that rewards a single successful run.
The platform half is not merely unmeasured; it is discarded, and defensibly. DeepSWE, an agentic coding benchmark that documents the decision in a dedicated section, excludes rollouts terminated by a model-provider error from both the numerator and the denominator of every metric, on the grounds that they “carry no signal about whether the agent could solve the task.” For measuring a model this is correct. The excluded fraction ranges from 0%, including the three top-ranked configurations, to 5.3% for Gemini 3 Flash [17]. Those runs terminated on a real platform, and once excluded no published figure retains them. No provider discloses its 429 or 529 rates. And no metric composes the two halves into the quantity a buyer requires: useful autonomous work per interruption, per dollar.
Each step is locally sound, but the binding term is absent.
What a Real Scorecard Would Measure#
None of this requires unusual instrumentation. A scorecard would run a standardized long-horizon workload and report mean time between interventions, disaggregated by cause: model-caused (drift, dead ends, derailment), access-caused (rate limits, overload errors, quota changes), and the category currently uninstrumented — runs that completed and should not have. It would publish 429 and 529 rates as service-level agreements publish uptime, and report quota variance alongside the quota itself. Because the third category cannot be observed from outside a single run, it would further require what no provider offers on a schedule: degradation postmortems carrying dates and blast radius, published on the same cycle as the uptime page. The objection that interruption rate is workload-dependent has real force. The same study finds reliability decay is domain-stratified, with graceful-degradation scores falling from 0.90 to 0.44 in software engineering while document processing barely moves, 0.74 to 0.71 [7]. But it proves too much. Every benchmark is workload-dependent, and standardization proceeds regardless.
I examined the public data for such a figure. An independent registry has polled every routing endpoint on OpenRouter at thirty-minute intervals since July, producing roughly three million observations across some four hundred models [18], [12]. Ranked on it, the models are indistinguishable. Across the 293 with a complete history the interquartile range of measured uptime is 0.12 percentage points, and the median is 99.98% in every price tier, from the cheapest models to those costing seventy-five times more. The rank correlation between price and availability is +0.09; price is the only capability proxy the public data makes machine-readable.
That is the smaller finding. The larger one is what the data cannot contain. Every state in it — up, degraded, down, idle — derives from a single binary question: whether the request returned. There is no field for a 429, none for a 529, none for a quota, a rate limit, or an error rate, and nothing describing the content of a response. Evaluated by this instrument, the August failure returns perfect health, correctly: requests succeeded, endpoints remained up, and the answers were wrong. The dataset the field treats as its reliability record measures the one failure mode that was never in question.
The instrument for the other mode does not exist. Until it does, “best model” will continue to answer the wrong question, and the answer will continue to be paid for in payroll.
References#
[1] “A postmortem of three recent issues,” Anthropic Engineering, Sep. 17, 2025. [Online]. Available: https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues. [Accessed: Aug. 29, 2026].
[2] “An update on recent Claude Code quality reports,” Anthropic Engineering, Apr. 23, 2026. [Online]. Available: https://www.anthropic.com/engineering/april-23-postmortem. [Accessed: Aug. 29, 2026].
[3] “Anthropic tightens usage limits for Claude Code — without telling users,” TechCrunch, Jul. 17, 2025. [Online]. Available: https://techcrunch.com/2025/07/17/anthropic-tightens-usage-limits-for-claude-code-without-telling-users/. [Accessed: Aug. 29, 2026].
[4] “429 resource_exhausted,” Google AI Developers Forum, thread opened Dec. 11, 2025; 50 posts, 17 participants, last activity Jun. 11, 2026. [Online]. Available: https://discuss.ai.google.dev/t/429-resource-exhausted/111737. [Accessed: Aug. 29, 2026].
[5] “Provisioned throughput units onboarding,” Microsoft Learn, Azure AI Foundry documentation. [Online]. Available: https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/provisioned-throughput-onboarding. [Accessed: Aug. 20, 2026].
[6] “Service tiers,” Anthropic API documentation. [Online]. Available: https://platform.claude.com/docs/en/api/service-tiers. [Accessed: Aug. 20, 2026].
[7] A. Khanal, Y. Tao, and J. Zhou, “Beyond pass@1: A reliability science framework for long-horizon LLM agents,” arXiv preprint arXiv:2603.29231, Mar. 31, 2026. [Online]. Available: https://arxiv.org/abs/2603.29231. [Accessed: Aug. 20, 2026].
[8] G. Orosz, “The Pulse: Did capacity shortages turn Anthropic hostile to devs?,” The Pragmatic Engineer, 2026. [Online]. Available: https://blog.pragmaticengineer.com/the-pulse-did-capacity-shortages-turn-anthropic-hostile-to-devs/. [Accessed: Aug. 20, 2026].
[9] “Anthropic unveils new rate limits to curb Claude Code power users,” TechCrunch, Jul. 28, 2025. [Online]. Available: https://techcrunch.com/2025/07/28/anthropic-unveils-new-rate-limits-to-curb-claude-code-power-users/. [Accessed: Aug. 29, 2026].
[10] “Models, usage, and limits in Claude Code,” Anthropic Help Center. [Online]. Available: https://support.claude.com/en/articles/14552983-models-usage-and-limits-in-claude-code. [Accessed: Aug. 29, 2026].
[11] “DeepSeek limits model access due to overwhelming server demand,” Engadget, Feb. 6, 2025. [Online]. Available: https://www.engadget.com/ai/deepseek-limits-model-access-due-to-overwhelming-server-demand-151339342.html. [Accessed: Aug. 20, 2026].
[12] dthinkr, “openrouter-uptime: independent uptime registry for every OpenRouter model,” GitHub repository. Half-hourly endpoint polls; derived/ series for Jul. 4 – Aug. 29, 2026 analysed here. [Online]. Available: https://github.com/dthinkr/openrouter-uptime. [Accessed: Aug. 29, 2026].
[13] “Time horizons,” METR, May 8, 2026. [Online]. Available: https://metr.org/time-horizons/. [Accessed: Aug. 20, 2026].
[14] T. Kwa et al., “Measuring AI ability to complete long software tasks,” arXiv preprint arXiv:2503.14499, Mar. 18, 2025. [Online]. Available: https://arxiv.org/abs/2503.14499. [Accessed: Aug. 20, 2026].
[15] S. Yao, N. Shinn, P. Razavi, and K. Narasimhan, “τ-bench: A benchmark for tool-agent-user interaction in real-world domains,” arXiv preprint arXiv:2406.12045, Jun. 2024. [Online]. Available: https://arxiv.org/abs/2406.12045. [Accessed: Aug. 20, 2026].
[16] “Claude Sonnet 4.6 system card,” Anthropic, 2026. [Online]. Available: https://anthropic.com/claude-sonnet-4-6-system-card. [Accessed: Aug. 20, 2026].
[17] W. Huang, C. Lee, L. Tng, and S. Ge, “DeepSWE: Measuring frontier coding agents on original, long-horizon engineering tasks,” arXiv preprint arXiv:2607.07946, Jul. 8, 2026. [Online]. Available: https://arxiv.org/pdf/2607.07946. [Accessed: Aug. 20, 2026].
[18] “How to evaluate LLM provider performance,” OpenRouter Blog. [Online]. Available: https://openrouter.ai/blog/insights/evaluate-llm-provider-performance/. [Accessed: Aug. 20, 2026].