In this article 6 sections

Anthropic has put unusually concrete operating numbers around its internal agent fleet. In August, roughly 30,000 agents were doing research and engineering work at any one time on its most-used internal platform. Anthropic says every action on that platform passed through an online monitor before execution, generating more than one billion decisions during the month. The monitor blocked 0.002%, or about one in 47,000. (Anthropic)

The blocked fraction is not a safety score. It does not show how many harmful actions the monitor missed, and Anthropic has only been operating this system for months. The disclosure is useful because it names three measurements that teams can actually run: coverage, review latency, and escalation rate. Reuters independently reported the headline figures but not an audit of the underlying system. (Reuters)

Oversight needs a service-level objective

Anthropic separates preventive and detective controls. Online monitors inspect 100% of actions before execution, usually within seconds. Humans review blocked actions within a week. Offline monitors ingest 100% of activity after execution, flag roughly one or two transcripts per thousand, and escalate about 50 high-priority cases per week to people after classifier triage.

That design turns oversight into an operating system rather than a policy statement. A principal engineer can ask whether every consequential action crosses the monitor, whether irreversible effects are reviewed before execution, whether the human queue stays within its latency target, and whether sampling non-alerts reveals misses. The answers belong on the same dashboard as availability and error budgets.

Conceptual flow showing every action on Anthropic's disclosed internal platform passing through an online monitor before execution. Allowed actions execute and later enter offline monitoring, while blocked actions and selected offline findings go to human review.
Conceptual map of the control split described by Anthropic, not an official system diagram. Online monitoring can stop an action before execution; offline monitoring reviews executed activity and sends selected cases into human escalation.

The architecture also preserves agent identity across model upgrades and links messages to source material and transcripts. This matters when many agents delegate to one another. Without durable identity and provenance, an investigator cannot distinguish a repeated claim from independent evidence or reconstruct which model, prompt, and tool policy produced an effect.

Anthropic also reports that Claude now leads 26% of its AI R&D work and collaborates on more than 90%, while no measured category is fully autonomous. Those figures come from a company-built automation index. Claude agents catalogued about 15,000 tasks, organized them into a 542-node tree, and another Claude judge assigned automation levels. Model and human raters agreed exactly 59% of the time, although they were within one level 97% of the time. The index is a serious measurement attempt, but still a self-measurement built partly by the system being measured.

Included Health kept domain ownership distributed

Included Health published a concrete counterexample to the idea that an agent platform must collapse domain logic into one general agent. Its Dot healthcare guide uses a main conversational router plus workflows owned by teams responsible for urgent care, scheduling, specialist search, behavioral health, and other services. Shared skills describe when a service applies, when it does not, and which edge cases require more questions. (LangChain case study)

Two implementation choices are worth copying. First, a global platform prompt carries voice and tone across agents while domain teams retain their own workflow logic. Second, a handoff includes both a conversation summary and a path to the full history, so the receiver can operate quickly without treating a lossy summary as the record.

Human escalation is part of the graph, not an exception outside it. The agent can pause, route a conversation to a care advocate, and later resume with the human exchange in context. Included Health says every conversation currently enters a clinical review queue. It reports clinician agreement above a 95% routing target, detection of more than 99% of high-risk situations in regular audits, and a 75% lift in chat engagement after an August launch.

Those are vendor-hosted, company-supplied results. The case study does not publish sample sizes, confidence intervals, the denominator for high-risk situations, or the false-positive burden. The architecture is therefore more reusable than the performance claims. For a healthcare deployment, the minimum evidence package should include sensitivity and specificity by risk class, review volume, handoff latency, disagreement categories, and the population used for each estimate.

Astra for Law is a configured system, not a new base model

OpenAI’s Astra for Law makes a similar systems point from a different domain. It combines GPT-6 Astra with legal instructions, a search index covering more than 230 million URLs, and firm-level governance controls. Free Law Project’s CourtListener supplies a large open collection of opinions and court records. (OpenAI) (Free Law Project)

On 200 questions from a private validation set of Vals AI’s Legal Research Bench, OpenAI reports 54.0% overall correctness at the highest reasoning effort, versus 38.7% for GPT-6 Astra with web search. On case-law questions it found 24% more reference cases. This is evidence that corpus design and retrieval policy can move end-to-end quality without changing the base model.

It is not yet evidence of better legal outcomes. The validation set is private, the provider ran the comparison, and early access is limited to selected firms. Teams evaluating a vertical system should separate authority retrieval, passage recall, citation validity, analysis quality, and workflow outcome. They should also test permissions and ethical walls as executable policy, not assume that zero data retention covers who may retrieve which matter.

Migration radar: Antigravity changed its tool contract

Google released antigravity-preview-09-2026 on September 17 and will shut down antigravity-preview-05-2026 on October 5. Remote-sandbox users that only consume text output can change the model string. Integrations that parse function calls or run tools locally need a real migration: built-in tool parameters changed from snake_case to PascalCase, and file edits moved from full rewrites to line-range replacements. (Gemini API changelog)

Treat that as a protocol version, not a model alias. Record the agent version with each trajectory, replay tool-call fixtures against both versions, and verify file mutations on disposable workspaces before switching production traffic.

Stay Sharp: pair fixed and rolling measurement baskets

Anthropic freezes its July task tree so month-to-month automation scores use the same denominator. That makes trend comparisons possible, but a fixed basket can miss work that appears, disappears, or moves to humans as the system changes.

Run two series. Keep a fixed basket for comparability, and rebuild a rolling basket to detect task substitution and new work. Publish both the score and the versioned task taxonomy. A rising fixed-basket score with flat or falling rolling-basket coverage is not broad automation; it may only mean the system is getting better at yesterday’s work.

What to watch

  • Anthropic plans independent embedded evaluators. The consequential evidence will be whether they can verify monitor misses, queue latency, and task classifications, not merely reproduce the published totals.
  • Google’s October 5 Antigravity shutdown is a hard migration boundary. Watch for compatibility tooling or a deadline change before assuming an alias swap is sufficient.
  • Astra for Law is coming to the API. Public benchmark items, retrieval traces, or independent firm evaluations would make the reported 54.0% easier to interpret.