In this article
OpenAI published a guide this week for connecting AI usage to business value. Its admin analytics can group a sample of messages into tasks, show where credits are spent, and track Codex contributions to merged code. More importantly, the guide tells teams to compare that activity with review time, defects, rework, downstream capacity, and financial outcomes. (OpenAI)
Anthropic published a different measurement system for a different purpose. It catalogued roughly 15,000 internal AI research and development tasks, organized them into a 542-node tree, and estimated how much work Claude assists, collaborates on, or leads. The company says about 30,000 agents were active at any one time on its main internal platform in August, generating more than one billion monitor decisions. (Anthropic) (Reuters)
These disclosures are not equivalent evidence. OpenAI is explaining its own enterprise product. Anthropic is measuring its own lab with a methodology partly executed and judged by Claude. Neither proves a general productivity gain.
They do expose the same architectural gap: most organizations can count AI activity, but cannot follow one unit of work from intent through execution, human intervention, operational effects, and business outcome.
The missing primitive is not another model dashboard. It is a task ledger.
The task is the join key
Token counts answer a capacity question. Active-user counts answer an adoption question. Neither says what work was attempted, whether it was completed, how much human effort it consumed, or what changed because of it.
OpenAI’s task classifier moves one step closer by grouping sampled messages into use cases and tasks. Its Codex view connects AI activity to engineering artifacts such as merged commits. The guide then recommends measuring review and correction time, quality, capacity released, and business value. That sequence matters because it separates several claims that are often collapsed into one:
- a person used AI;
- the system produced an artifact;
- the artifact passed review;
- the workflow completed;
- the completion changed an outcome.
Each transition can fail independently.
Anthropic took a more elaborate route. It sampled work records, generated a task catalogue, froze a task tree for longitudinal comparison, assigned automation levels, and weighted categories by estimated person-time. Its published result says Claude leads 26% of measured AI R&D work and collaborates or leads on more than 90%. No measured category is fully autonomous. (Anthropic)
The useful part is not the headline percentage. It is the explicit unit of analysis. Anthropic had to decide what a task is, version the taxonomy, preserve a denominator, and state how ratings are produced. It also reports the uncomfortable parts: the model judge and human raters agreed exactly 59% of the time, although they were within one level 97% of the time; a frozen basket can miss new work; and the system being measured helps classify the measurement.
That is what a serious task ledger needs to preserve:
- Task identity: a stable work item that survives model, prompt, and tool changes.
- Taxonomy version: the definition used when the task was classified.
- Execution lineage: model, harness, tools, data sources, policy version, and retries.
- Human involvement: instruction, review, correction, escalation, and approval time.
- Operational result: completed, failed, abandoned, duplicated, or recovered.
- Outcome: quality, cycle time, cost, revenue, risk, or another domain measure.
- Evidence strength: observed, self-reported, model-classified, or independently verified.
Do not force all domain state into an observability vendor. Keep authoritative workflow records in the systems that own them, then emit a shared task identifier and versioned events into the analytical plane. The ledger is a join model, not a new source of truth.
Outcome denominators decide whether an application is real
This week’s application reports show why the denominator matters.
Salesforce says Siemens uses agents for roughly 2,500 in-scope inbound leads per month across 132 countries. A model conducts the conversation and writes structured qualification fields; about 50 predefined rules decide how a lead is routed. Salesforce reports an 11% engagement rate, a 6% qualification rate, and positive ratings for 80% of interactions. The account does not publish the measurement window, rating instrument, comparison design, or the false-positive cost of sending the wrong lead to a seller. (Salesforce)
LangChain’s Included Health case study reports clinician agreement above a 95% routing target, detection of more than 99% of high-risk situations in regular audits, and a 75% lift in chat engagement after an August launch. It does not provide sample sizes, confidence intervals, the denominator for high-risk cases, or the review burden created by false positives. The architectural details are still useful: domain teams own separate healthcare workflows, handoffs carry both a summary and a path to full history, and every conversation currently enters clinical review. (LangChain)
These are stronger than a demo and weaker than controlled outcome evidence. They show deployed workflows and expose some operational numbers. They do not establish how much value came from the model, the surrounding process change, additional review, or selection of easier cases.
The wider research makes a universal productivity multiplier even less plausible.
A randomized field study across Microsoft, Accenture, and another large company reported 26.08% more completed tasks among 4,867 developers given access to a coding assistant, with larger gains among less experienced developers. (Field experiments) A separate randomized study of experienced open-source developers working in familiar repositories found that early-2025 tools made participants take 19% longer. (METR RCT)
METR later said its follow-up experiment could not provide a reliable current estimate. Developers increasingly declined to participate when assigned to work without AI, and parallel use of multiple agents made time measurement harder. (METR update) In another 2026 survey, 349 technical workers reported a median 1.4 to 2 times increase in the value of their work, while METR cautioned that the magnitude should be treated skeptically. (METR survey)
Those results are not a contradiction to average away. They describe different workers, tasks, tools, periods, and measurement methods. The decision-relevant question is not “Does AI improve productivity?” It is:
For this versioned workflow and task population, does AI improve the outcome after review, rework, failures, and displaced human effort are counted?
A platform team should therefore reject portfolio metrics that cannot be decomposed by task class. Aggregate adoption can rise while a high-value workflow deteriorates. Completion can rise while defect repair consumes the gain. A faster first draft can lengthen approval. Capacity can be released without being converted into revenue, throughput, or service quality.
Incidents belong in the same ledger
A task ledger is also an operational control surface.
OpenAI’s managed Agents API had an incident on September 14 in which customers experienced delays or could not start turns in managed sessions. The service later recovered. (OpenAI Status) That event is not captured by model accuracy, token spend, or task classification. It appears in the transition between accepted work and executable work.
The weekend’s more serious example came from a cybersecurity evaluation. Reuters reported that a Gemini model accessed three real companies after evaluator Irregular gave the system internet access and it treated real targets as in scope. Google said the model stopped in each case, the affected entities were notified, and the testing process was changed. Irregular said similar scope problems had affected tests of other labs’ systems. (Reuters)
The architectural lesson is narrower than “agents are dangerous.” Scope is a runtime dependency. A work item needs a machine-checkable record of:
- the intended target and allowed environment;
- the evaluator or user who established that scope;
- the credentials, network paths, and tools made available;
- the external effects actually produced;
- the signals that stopped, redirected, or escalated the work.
Without that lineage, an incident review cannot distinguish a model choosing an unauthorized action from a harness granting access to the wrong environment, a stale scope definition, or a credential boundary that failed open.
The same record supports value measurement and incident response. Both need to know what was requested, what executed, which human and automated controls intervened, and what outcome followed. Maintaining separate “ROI telemetry” and “safety telemetry” creates two partial stories about the same work.
Build the ledger before the executive dashboard
The practical sequence for a Principal or Architect is small enough to start with one consequential workflow.
Choose a bounded task population. Pick work with an authoritative start, terminal state, and domain owner. Lead qualification, care routing, and code changes can qualify. “Knowledge work” cannot.
Define terminal outcomes before instrumenting prompts. A lead is not complete when the agent stops speaking. A code task is not complete when a patch exists. Record the accepted business state, including rejection, escalation, and abandonment.
Carry one task identifier across boundaries. Propagate it through model calls, tool invocations, queues, human review, and the system of record. Preserve parent-child links when an agent delegates work.
Version every classifier. Task categories, risk labels, quality graders, and automation levels are models too. Store their version, confidence, and evidence. Reclassify historical samples when a taxonomy changes instead of silently splicing incompatible series.
Price human intervention. Review time, corrections, escalations, and exception handling belong in cost per completed task. A workflow that saves generation time but creates a noisy review queue may move work rather than remove it.
Keep an exposure log. Record external writes, messages, file publication, credential use, and target systems as effects attached to the task. This is the bridge between business observability and security investigation.
Only after those records are stable should the organization aggregate them into portfolio dashboards. Otherwise a polished chart hides shifting denominators and incomplete work.
Stay Sharp: activity is not causal impact
A task ledger improves observability. It does not by itself prove that AI caused an outcome.
Users select themselves into AI tools. Teams assign AI to some tasks and not others. Managers change workflows, training, and staffing during rollout. Models improve while the task mix changes. These effects make a before-and-after chart easy to produce and hard to interpret.
The strongest practical design is a randomized rollout. Randomize access or treatment at the user, team, queue, or task level, then measure the same terminal outcomes for both groups. Use cluster randomization when people share work and treatment would spill across individuals. Define the analysis and exclusion rules before reading the results.
Randomization is not always possible. Three weaker designs can still improve on a simple trend line:
- Staggered rollout: compare early and later groups over the same calendar period, controlling for team and task mix.
- Matched cohorts: pair treated tasks with similar untreated tasks using pre-treatment attributes, then test sensitivity to unmatched differences.
- Interrupted time series: model the pre-rollout trend and look for a durable level or slope change, while recording other changes that occurred at the boundary.
Every design needs guardrail outcomes. If throughput rises, track defects, rework, customer complaints, review latency, and security events. If AI shifts work to reviewers, report total labor, not just authoring time. If a team handles more easy tasks and fewer hard ones, publish the task-mix change beside the aggregate.
Four questions belong in the design review:
- What is the unit of assignment: person, task, queue, or team?
- What is the terminal outcome, and who owns its truth?
- Which costs and harms could move outside the measured boundary?
- What result would cause us to reduce or stop the deployment?
The final question is the most neglected. A measurement system built only to justify expansion is not evaluation infrastructure. It is sales instrumentation.
This week did not produce a trustworthy universal AI productivity number. It produced something more useful for architecture: a clearer shape for the records that would make workflow-specific claims testable. Count tasks, preserve lineage, price intervention, record effects, and evaluate outcomes against a credible counterfactual. Tokens can remain a capacity metric. They should not be mistaken for value.