In this article 8 sections

The most consequential development in the last day is not a new base model. It is the agent harness becoming a product boundary.

OpenAI has exposed the runtime behind Codex through a public Agents API: durable sessions, managed orchestration, automatic context compaction, tool discovery, subagents and optional hosted sandboxes. On the same day, Anthropic added an auto permission mode to Claude Managed Agents that evaluates each built-in or MCP tool call server-side and decides whether to run it, deny it or pause for human approval. (OpenAI) (Anthropic release notes)

Taken together, these releases make a platform shift explicit. The unit enterprises will increasingly buy is not just model + API. It is model + state machine + context policy + tools + sandbox + authorization + observability. That can remove a great deal of brittle application code. It can also move more of your operational semantics into a provider-specific control plane.

A separate Anthropic research release makes the stakes clearer. Its new military and intelligence evaluations show models performing useful target-identification, geolocation and simulated weapons-engineering work, including meaningful capability from a model the study describes as open-weight and below the frontier. The exact benchmark results are provider-run and should not be treated as independent rankings, but the architectural conclusion does not require a leaderboard: dangerous capability is becoming a systems problem before it is exclusively a frontier-model problem. (Anthropic research)

Operational priority: Datasette’s September 11 security announcement urges operators of public instances, particularly those protecting private data with authentication plugins, to upgrade to the patched release for their series: 1.0a39 or 0.65.4. The maintainers report that AI-assisted auditing helped uncover additional issues. This is an immediate patching action; it is not evidence that an automated audit alone establishes security. (Datasette)

OpenAI is turning the Codex harness into managed infrastructure

OpenAI launched the Agents API in public beta on September 10. The API exposes the harness and infrastructure behind Codex while letting developers choose where agent work executes: an OpenAI-hosted sandbox, their own infrastructure or a supported sandbox partner. OpenAI says there is no separate Agents API fee during the beta; customers pay for the tokens and tools the agent consumes. (OpenAI)

The more important details are in the runtime semantics. OpenAI describes durable sessions for long-running work, managed session orchestration and recovery, automatic context compaction, dynamic tool search that loads tool definitions only when needed, programmatic tool calling and subagents with isolated contexts. The underlying Codex harness is publicly inspectable as an open-source codebase, while OpenAI operates and evolves the managed version. (OpenAI)

This is a useful dividing line for platform architecture. For years, an “agent framework” often meant application code implementing a loop around a model call. The managed-runtime abstraction absorbs several difficult distributed-systems concerns:

  • preserving state across hours or days;
  • deciding which context survives a long trajectory;
  • keeping tool catalogs from consuming the working context;
  • scheduling parallel subordinate work;
  • recovering after process or environment failure;
  • and maintaining an execution environment with files and code.

These managed services can remove substantial operational work. Their semantics remain provider behavior you must evaluate.

A model portability layer that normalizes only messages, tools and responses is therefore becoming insufficient. If an application depends on a provider’s compaction algorithm, session persistence, subagent topology, sandbox lifecycle or tool-discovery semantics, migration requires reproducing the harness contract, not merely changing the model identifier.

I would make that coupling explicit in architecture records. Decide which parts of the agent runtime are acceptable managed dependencies and which need portable representations: authoritative task state, artifact storage, tool contracts, permission policy, trace export and the evaluations that define acceptable behavior.

Anthropic is moving tool permission decisions into the server-side control plane

Anthropic’s September 10 Claude Platform release adds a third permission policy for Managed Agents. Alongside always_allow and always_ask, auto lets Anthropic’s server evaluate every agent-tool or MCP-tool invocation and then run, deny or pause the call for approval. Tool-use events expose both an evaluation and an evaluated_permission, giving the runtime a record of how the call was classified. (Anthropic release notes)

That is a notable step beyond a static confirmation switch. It turns permission handling into a policy-decision point inside the managed agent runtime.

The distinction between availability and authorization still matters. Anthropic’s documentation is explicit that permission policies govern whether enabled server-executed tools run; disabling a tool removes it from the agent entirely, while custom tools execute in the customer’s application and remain the customer’s responsibility. MCP tools also default more conservatively than the built-in agent toolset. (Anthropic permissions)

For an enterprise platform, I would model this as three separate layers:

  1. Capability exposure, is the tool present at all?
  2. Policy decision, may this invocation proceed automatically, be denied or require approval?
  3. Effect enforcement, do the downstream identity, network and resource policies actually permit the operation?

The new auto layer can reduce approval fatigue, but it should not become the sole security boundary. A server-evaluated permission decision should sit alongside independently enforced constraints such as scoped credentials, egress restrictions, tenant authorization and transaction limits. Yesterday’s Daily focused on why model-generated reasoning is weak independent security evidence; today’s release shows the complementary design direction: make the permission decision itself a first-class runtime event that can be logged, evaluated and constrained by infrastructure outside the conversation.

One operational constraint deserves explicit treatment: running sessions retain the tool configuration they started with. Updating an agent’s permissions affects newly created sessions. A policy rollout therefore needs a plan for existing sessions, not just an updated agent definition. (Anthropic permissions)

Military-domain capability is diffusing below the absolute frontier

Anthropic’s Frontier Red Team published new evaluations on September 10 covering intelligence targeting and conventional-weapons engineering. The work includes simulated cross-platform identity correlation, image and text geolocation, drone terminal-guidance software, payload delivery and navigation under GPS denial or spoofing. (Anthropic research)

There are important limitations. Several datasets are synthetic or simplified; the weapons work uses simulation rather than hardware; Anthropic ran the evaluations itself; and the study does not directly measure real-world uplift. Anthropic explicitly presents several results as suggestive rather than absolute performance estimates. (Anthropic research)

Even with those caveats, one result is strategically important: useful capability is not confined to the single strongest closed model. Anthropic reports that Kimi K3, described as open-weight in its test set, was behind the frontier overall but competitive with stronger models on some targeting tasks and capable of non-trivial performance on several simulated weapons-engineering tasks. The gap widened on the hardest settings, but it did not reduce the weaker model to irrelevance. (Anthropic research)

For architects and security teams, that changes the threat model. Governance that assumes only a handful of frontier endpoints need strong controls will age poorly. Capabilities diffuse into cheaper hosted models, open-weight checkpoints and locally deployable systems. Controls therefore need to attach to workflows, tools, data and effects, not only to the provider’s model tier.

This also complicates release-policy debates. A capability threshold can still matter, but the operational question is increasingly continuous: what does this model become capable of when combined with search, private data, iterative simulation, code execution and a persistent agent loop?

Infrastructure radar: custom inference silicon is joining NVIDIA’s rack, not escaping it

NVIDIA and d-Matrix announced that d-Matrix’s next-generation Raptor inference XPUs will integrate with NVIDIA’s NVLink Fusion, MGX rack architecture and Spectrum-X networking. NVIDIA positions the arrangement as a way for custom accelerators to reuse a mature rack, networking, power, cooling and supply-chain platform rather than building an independent data-center stack. d-Matrix targets initial availability in Q4 2027 and expects Raptor’s tape-out, the handoff of its chip design for fabrication, before the end of 2026. These are roadmap targets, not shipping capacity. (NVIDIA) (d-Matrix)

The strategic signal is subtle: accelerator competition does not necessarily weaken NVIDIA’s infrastructure position. Specialized inference silicon can diversify the compute layer while standardizing more of the surrounding system on NVIDIA interconnect and rack architecture. For platform teams, hardware portability will increasingly depend on runtime and workload placement across heterogeneous accelerators, not on pretending the physical infrastructure is vendor-neutral.

Engineering radar: Shopify is revisiting the cost of sharing code

Shopify’s September 10 engineering report says it is moving its mobile apps from React Native to Swift and Kotlin because coding agents have reduced the cost of implementing and maintaining features twice. The team describes small review checkpoints, behavior tests, visual inspection and human approval, alongside headless business logic that agents can exercise quickly. (Shopify Engineering)

This is a company-specific experience, not a general case for rewriting cross-platform apps. Its useful lesson is to revisit architecture decisions when their cost assumptions change, and measure verification, maintenance and platform parity alongside code-generation speed.

Stay Sharp: context compaction is lossy compilation, not summarization

Automatic context compaction sounds like a convenience feature, but in a long-running agent it is part of the system’s state-management semantics.

A transcript contains several kinds of information with very different durability requirements: current instructions, decisions, failed approaches, tool observations, external artifact locations, unresolved questions and evidence supporting earlier conclusions. Compressing all of that into a prose summary is equivalent to compiling a rich execution history into a smaller representation. The compiler can introduce semantic loss.

A robust long-running agent should therefore separate at least three layers:

authoritative state, versioned files, database records, issue state, test results, external resource IDs and other facts that should survive independently of the model context;

trajectory evidence, an append-only or reconstructable event history of messages, tool calls, results, approvals and important state transitions;

working context, the compact representation presented to the model for the next decision.

Compaction should transform the third layer, not silently replace the first two.

The design test is simple: after several compaction cycles, could a fresh process reconstruct why the current task state is valid without trusting the compacted narrative? If not, the summary has become an accidental source of truth.

Useful compaction pipelines preserve explicit invariants: task goal, hard constraints, accepted decisions, unresolved blockers, artifact identifiers and pointers back to evidence. They can discard conversational redundancy and stale exploration while retaining a route to the underlying history. Evaluations should include long-horizon cases where an early constraint or failed experiment becomes relevant much later; otherwise a compactor can look excellent on token reduction while quietly degrading correctness.

OpenAI making automatic compaction part of a managed agent runtime is therefore not merely a context-window optimization. It is a reminder that memory policy is executable architecture.

Watchlist

  • Agents API portability: watch whether OpenAI publishes stronger export/replay semantics for durable sessions and compaction state as the beta evolves.
  • Managed permission evaluation: watch how Anthropic exposes policy rationale, calibration and enterprise override controls around auto, and whether comparable policy-decision primitives appear across providers.
  • Independent replication of military-domain evals: provider-run results are enough to flag the risk direction, not enough to settle comparative model capability.
  • September dates: OpenAI lists September 24 as the Sora 2/Videos API shutdown date. (OpenAI deprecations) Anthropic separately targets September 30 for Phase 1 of its frontier-safety “Moonshot R&D” projects; this is a company-set research goal that can change. (Anthropic roadmap) Track the service shutdown and the research commitment separately.

Further reading