In this article 6 sections

Anthropic changed the economics of a terminal state on September 24. A refusal that arrives before any output is now billable when stop_details.category is bio, frontier_llm, or reasoning_extraction. The request is charged at the normal rate of the model that ran. Pre-output refusals in other categories remain unbilled, and mid-stream refusals were already billable. (Claude Platform release notes)

The operational detail matters more than the invoice. Claude returns a classifier refusal as HTTP 200 with stop_reason: "refusal", an empty content array when no output was produced, and token counts in usage. In other words, the API call succeeded, compute was consumed, and the user task did not complete. (Anthropic refusal documentation)

That combination breaks dashboards that equate a 2xx response with useful work. It also makes an unbounded retry loop a cost amplifier.

HTTP success is not task success

A production caller now needs at least three independent results for every model attempt:

  1. Transport: did the API accept and complete the request?
  2. Model outcome: did it answer, refuse, time out, or fail partway through a stream?
  3. Product outcome: did the workflow deliver the requested result, fall back, escalate, or stop safely?

The refusal category belongs in the second ledger. It is machine-readable and determines billing behavior today, but Anthropic says the billed set may change as it measures false positives. The human-readable explanation is explicitly unstable and should be displayed, not parsed.

This distinction should reach both metrics and control flow. Track refusal rate by requested model, category, route, customer workflow, and prompt version. Count the refused attempt in cost and rate-limit consumption even if the final user response comes from a fallback. Alert on changes in category mix, not only aggregate HTTP error rate.

For capacity planning, use an attempt-level record rather than a request-level total:

requested model -> refusal category -> fallback model -> final outcome

That record lets finance reconcile charges, safety teams investigate classifier shifts, and reliability teams see whether fallback capacity is absorbing or merely hiding refusals.

Conceptual attempt-ledger diagram. A requested model returns an HTTP 200 refusal, a fallback model returns an answer, and separate rows record transport success, attempt-level cost and rate limits, and the recovered product outcome.
Conceptual attempt ledger synthesized from Anthropic’s refusal and fallback documentation, not an official protocol diagram. Server-side fallback is triggered by a classifier refusal, not by rate limits, overloads, or server errors. Side effects still require a separate idempotency record.

Fallback creates two ledgers

Anthropic’s server-side fallback can retry a refusal within one API call. The final response names the model that served it, while usage.iterations records each attempt. A billed refusal and the fallback request can therefore both contribute to cost. Anthropic’s fallback credit compensates for the fallback request’s prompt-cache miss, but it does not erase the refused attempt. Every attempted model also consumes its own rate limit. (Anthropic refusal documentation)

The top-level response is not a complete cost record because its usage covers only the model that returned the message. Cost accounting must read usage.iterations. Reliability accounting must also keep the distinction between a recovered refusal and a first-pass answer. Otherwise a healthy fallback masks a deteriorating primary route until fallback latency, capacity, or policy behavior changes.

Retries need a budget expressed in attempts, money, and elapsed time. They also need an idempotency policy for tools. A model fallback is safe only if the workflow knows which external effects already happened before the refusal. Anthropic says completed server-tool results can carry into a fallback, while an open tool-use block in a streaming response can stop the retry path. That is an execution-state boundary, not a prompt detail.

Identity must scope every stateful layer

LangChain’s same-day Managed Deep Agents 0.8 release adds a useful second signal. It separates durable agent memory from user memory, keys user memory to the authenticated caller, and says the runtime does not copy content between those layers. Its defaults allow user memory in one-to-one Slack messages but deny it in Slack channels, group messages, and HTTP channels. (LangChain)

Credentials receive the same split. Agent-owned credentials are shared across users, while user-owned credentials carry an individual’s permissions into services such as GitHub, Linear, or Notion. This is the right shape for a multi-user agent, but the vendor post is design documentation, not an independent security assessment.

The architectural lesson is that identity cannot be attached only at login. It must parameterize memory namespace, credential selection, channel policy, tool authorization, trace access, and deletion. A request that moves from a direct message to a group channel has changed authority even if the prompt text is identical.

Review that boundary with negative tests. A group thread should not retrieve direct-message memory. An HTTP webhook without a trusted user mapping should not inherit a person’s credentials. A revoked identity should lose both future tool access and access to user-scoped memory. The default-deny choices matter most where the channel cannot prove a single caller.

Local execution is another route, not an escape hatch

Google’s September 23 Antigravity SDK update adds local model execution through LiteRT and an OpenAI-compatible server interface. Its published hybrid demo uses a cloud model to plan from filenames and task descriptions, then sends code auditing and patching to local Gemma workers. Google reports that 3,322 tokens, or 97.2 percent of the recorded run, stayed local. (Google Developers Blog)

That is a vendor-run demonstration, not a comparative benchmark. It still illustrates a useful routing class. Local execution can change data exposure, marginal API cost, and offline behavior without changing the agent harness. It does not remove governance. The local model needs its own version, policy, resource ceiling, evaluation set, and terminal-state metrics. The example also recommends more than 24 GB of VRAM or unified memory, so local is a capacity decision as well as a privacy decision.

The routing ledger should therefore record where an attempt ran, which data crossed the boundary, and why the route was selected. “Local” is not synonymous with “safe,” just as HTTP 200 is not synonymous with “successful.”

Stay Sharp: count terminal states before retries

Use a small terminal-state model for every agent step:

  • answered: useful result returned by the requested route
  • recovered: another route completed after a refusal or failure
  • refused: policy terminal state with no useful result
  • failed: infrastructure or tool failure
  • aborted: budget, deadline, or operator stopped the work

Attach cost, latency, model, route, and side-effect status to each attempt. Then define the product SLO on useful outcomes, not HTTP responses. A refusal can be correct safety behavior and still consume budget. A fallback can recover the task and still reveal a primary-route regression. Both facts need to survive aggregation.

What to watch

  • Whether Anthropic changes the set of billed refusal categories or publishes category-specific false-positive evidence.
  • Whether fallback-heavy applications expose attempt-level cost and rate-limit data instead of only the final response.
  • Whether identity-scoped agent platforms publish adversarial tests for channel changes, memory isolation, credential revocation, and deletion.
  • Whether local-agent frameworks report reproducible quality, resource, and failure measurements alongside token-locality claims.