In this article 6 sections

OpenAI and Anthropic released new frontier models on September 22, but the more consequential engineering change sits below the benchmark tables. OpenAI’s GPT-6 prompt cache now reports when a changed tool definition breaks reuse. Anthropic’s Claude Opus 5.5 can reject request shapes that worked with Opus 5, and a model switch can discard preserved thinking from an existing conversation. (OpenAI caching) (Anthropic migration guide)

Together, those changes make one architectural point concrete: an agent’s tool catalog, conversation prefix, and model choice are runtime state. Updating them is closer to a state migration than a stateless endpoint swap.

OpenAI made cache misses observable

GPT-6 can reuse shared prompt prefixes within a 30-minute window. OpenAI says cached input receives a discount of up to 90%. The company also added a dashboard and a diagnostic tool that attributes misses to causes such as tools_changed. One published example shows 5,629 tokens affected by that reason. (OpenAI caching)

The diagnostic is more useful than the discount headline. A team can now distinguish an organic workload shift from a deployment that accidentally destroys its cache hit rate.

OpenAI recommends keeping tool definitions, schemas, and ordering stable. If a tool should be unavailable for one turn, use allowed_tools or tool_choice: none instead of removing it from the catalog. Change reasoning effort through configuration_update, and append new developer messages rather than editing earlier context. The goal is to preserve the longest possible shared prefix.

That guidance changes how feature flags should work. A flag that reconstructs the tool array is not only a capability decision. It can create a cost and latency event across every active conversation. Tool availability and tool identity therefore need separate representations.

The evidence is still provider-operated. OpenAI’s launch materials include customer-reported gains, but they do not provide an independent measurement of hit rates across diverse production workloads. The actionable part is the newly visible cache reason and the request controls, which teams can test against their own traffic. GPT-6 Sol and Luna are now available through the API at $2 and $0.10 per million input tokens respectively, before cache discounts. (OpenAI launch)

Claude Opus 5.5 turns an alias change into a protocol migration

Anthropic’s Opus 5.5 migration guide names four breaking changes. Adaptive thinking is always on, so requests that disable thinking or set a manual token budget return HTTP 400. Forced tool selection with tool_choice: any or a specific tool also returns 400. Anthropic directs applications toward automatic selection plus strict tool use or structured outputs. (Anthropic migration guide)

Conversation state has a second compatibility boundary. Preserved thinking blocks are bound to the model and the conversation prefix that produced them. Switching models may silently drop those blocks. Editing earlier system instructions, tools, or messages can instead produce a 400 for newer accounts. Anthropic recommends append-only conversations and mid-conversation system messages when the policy needs to change.

The computer-use surface differs by provider too. computer_20251124 is rejected on the Claude API and Google Cloud for Opus 5.5, which require computer_toolset_20260801, while the older tool can still run through Amazon Bedrock. The same model name therefore does not imply the same tool protocol on every host.

There is also an observability trap. Text that previously appeared between tool calls can move into thinking blocks, so a progress display may go quiet even while the agent continues working. Anthropic’s release notes add mid-conversation tool definitions as a beta feature that can add or remove schemas without editing the initial tool set or invalidating the prompt cache. (Anthropic release notes)

A router needs a compatibility plan

Price and benchmark quality are insufficient routing inputs for a long-running agent. Add a compatibility record for each supported route:

  1. Pin the provider, model ID, request schema, thinking mode, tool protocol, and response-shape assumptions.
  2. Replay representative conversations that contain tool calls, preserved reasoning, refusals, and partial failures.
  3. Record cache writes, cache reads, miss reasons, first-token latency, and the model that actually served the request.
  4. Treat HTTP 400, an HTTP 200 refusal, and a provider fallback as different operational events.
  5. Rehydrate a conversation from durable business state when hidden provider state cannot move safely.
Conceptual agent-route migration map. A versioned route contract contains the provider and model, request schema, thinking mode, tool protocol, and ordered prefix. Changing those fields can cause request rejection or fallback, a cache miss with cost and latency, or loss of preserved reasoning. Durable task status, approvals, tool results, identifiers, and effects remain outside the transcript so the session can be rehydrated safely.
Conceptual migration map synthesized from OpenAI’s cache guidance and Anthropic’s migration guide, not an official provider protocol. Version the route contract, replay it, observe the served path and cache reason, and rehydrate from durable state when provider-held continuity cannot move safely.

Canary tests should change one component at a time. A clean answer from a fresh prompt does not prove that an hour-old agent session can survive the migration. The replay set needs the same tool catalog and conversation prefixes that production uses.

The durable state should remain outside the model transcript: task status, approvals, tool results, external identifiers, and irreversible effects. Preserved reasoning and cached prefixes are accelerators, not the system of record.

Technical reading: faster parsing is not a routing policy

LlamaIndex published a September update to LiteParse with vendor-measured reductions of 20% to 25% in PDFium extraction time. It reports 2.76 milliseconds per page for text extraction and 3.94 milliseconds for Markdown on its test set, plus higher table scores and a new is-complex endpoint for routing difficult documents. (LlamaIndex)

The routing idea is useful, but it is close to the selective-fidelity OCR pattern covered on September 15. The new measurements do not establish how often the complexity classifier misroutes difficult pages, so this edition leaves it as technical reading rather than another architectural conclusion.

Stay Sharp: cache keys are deployment contracts

A prompt cache key captures more than text. It encodes the ordered prefix, tool schemas, model configuration, and sometimes provider-specific state. That makes it a deployment contract with three failure modes: a correctness break, a cost regression, or a continuity loss.

Version that contract. Before rollout, estimate which active sessions will retain their prefix, which will need a clean restart, and which request fields will become invalid. After rollout, compare cache hit rate and miss reasons by application version. If a tool-policy change is meant to be reversible, prove that the rollback restores both behavior and cache reuse.

What to watch

  • OpenAI’s Videos API and Sora 2 are scheduled to shut down on September 24 with no replacement listed. Legacy model shutdowns follow on September 28. (OpenAI deprecations)
  • Anthropic’s inline tool-definition beta needs production evidence on versioning, replay, and provider parity before teams treat it as a portable conversation feature.
  • Cache dashboards will be most valuable when teams connect miss reasons to deploys and feature flags, rather than reading a global hit-rate average.