In this article 6 sections

Microsoft and Hugging Face released a benchmark on October 3 that runs 507 stateful business workflows twenty times each. One result should change how teams read agent leaderboards. Kimi-K3 completed 93.89% of the tasks at least once, yet completed only 68 of 507 tasks in all twenty recorded attempts. That is 13.41% of the suite. Claude Opus 5 covered fewer tasks at least once, but completed 173 more tasks consistently. (Microsoft and Hugging Face)

This is not a contest between two models. It is a warning about two different product claims. “The agent can do this” is a capability claim. “The service will do this correctly when customers depend on it” is a reliability claim. Best-of-many success can support the first while concealing the second.

Three other pieces of work sharpened where that reliability must come from. A scientific workflow showed that selecting the right problem can matter more than extending a general agent. A training paper moved some harness behavior into model weights but left domain capability outside. A customer-support system retrieved policies, connectors, and procedures under three different failure contracts. Together they argue against treating the model, harness, and task definition as interchangeable sources of quality.

Twenty clean runs expose a different leaderboard

ThinkingBox starts every attempt from a fresh database and evaluates the terminal backend state. Most of its 507 synthetic workflows are checked only through deterministic state assertions. Thirty also evaluate required properties of the final response. This matters because plausible text and valid tool calls are not the outcome. (paper)

The common-set analysis covered 121,680 valid trials across twelve models. Of 79,853 attempts that failed executable checks, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. Among those failures, 77.61% contained wrong field values, 43.30% created unintended extra effects, and 25.36% omitted required effects. The categories overlap. (Microsoft and Hugging Face)

Last week’s review separated an authorization receipt from a task ledger. ThinkingBox adds a new and narrower result: even a well-defined terminal state does not make performance stable. A task ledger can tell you which attempts reached the right state. It cannot turn a model that succeeds intermittently into a dependable service.

The deployment gate should therefore contain three numbers, not one:

  1. Attempt success: the probability that one attempt reaches the required terminal state without extra effects.
  2. Repeat consistency: the share of tasks that meet a declared success threshold across independent resets.
  3. Operational error rate: failures caused by the evaluation system, tool service, timeout, or provider, reported separately from model outcomes.

ThinkingBox counts unresolved system errors as unsuccessful trials. That is conservative for a model comparison, but appropriate for a customer-facing service. The customer consumes the whole path. An internal diagnostic may separate model and infrastructure faults, while the service-level indicator must include both.

The benchmark still has limits. Its businesses, customers, and policies are synthetic reconstructions. Repeating the same task twenty times measures stability under sampled model behavior, not resilience to policy drift, adversarial input, changing data, or real user recovery. It is a better reliability test than pass@1, not a production reliability certificate.

A production case shows where the missing reliability can come from

An undated case study from agent vendor Context was publicly available by the October 4 cutoff. It reports a deployment across Qualcomm’s customer engineering organization that grew from one pilot team and five workflows to 85 teams and 1,600 production workflows over 17 weeks. Context says the same foundation model improved from a 23% to a 98% pass rate on a held-out suite of production-shaped support tasks. (Context and Qualcomm)

“Same model” does not mean “same system.” The deployment added indexed documentation, repository access, learned memories, retrieval tuning, harness fixes, and a failure-driven evaluation loop. The page says model weights and decoding settings stayed fixed, and it describes a context ablation that also held the harness fixed. Across the full 17-week rollout, however, harness changes were one of several shipped interventions.

The account is stronger than a launch testimonial because it names the evaluation window, strict task-specific rubrics, three grader passes per task, and which cost claim is modeled rather than measured. It is still commercial evidence from the vendor operating the system. The public page does not disclose the held-out task count, grader agreement, per-family failure distribution, human correction burden, or an independent audit. Its reported 40% productivity improvement is measured against a pre-deployment baseline rather than a randomized control.

This case complements ThinkingBox without resolving it. It shows that institutional context and an eval-gated improvement loop can close many failures without changing model weights. It does not show that a 98% held-out pass rate will repeat across independent live runs. A production reliability review should therefore keep four results separate: the context ablation, the total system improvement, repeated consistency by task family, and live outcomes after human correction.

Task selection is part of the system design

Matthew Schwartz’s October 1 account of BootLoops provides a different application lesson. BootLoops is an open-source, model-agnostic harness for exact calculations in quantitative science. Schwartz reports work across eighteen fields, much of it built by finding problems whose inputs, transformations, and checks suit current agents. He calls them “Claude-shaped” problems. (Anthropic)

The important detail is what happened after a technically correct result. In ecology and population genetics, outside experts were often unimpressed by the initial question. They redirected the workflow toward a residual worth explaining or a correlation that the field had not adequately measured. The computational result was not enough. Scientific taste entered through problem selection and interpretation.

One project has a more inspectable public record. An NBER working paper describes a workflow applied to 4,452 replication packages from five economics journals. It flags discrepancies in 3,460 articles or appendices, finds calculations it can accelerate by more than ten times in 496 articles, and proposes aligned extensions in 923. Those are workflow outputs, not a claim that every discrepancy invalidates a paper or every extension is scientifically important. The paper is a working paper rather than peer-reviewed evidence. (NBER)

Schwartz also discloses that he was a visiting researcher at Anthropic during the project. BootLoops is not an Anthropic project, but the surrounding account is still provider-hosted and many reported scientific results remain under verification.

The architecture lesson survives those caveats. Task discovery deserves its own pipeline:

  • define which inputs are available and legally usable;
  • identify transformations the agent can execute and checks it cannot easily game;
  • ask a domain expert whether the result would change a decision or research program;
  • only then invest in scaling the agent loop.

This is more than “keep a human in the loop.” The expert is not merely approving an answer. The expert changes the objective. Teams that automate a poor question faster have improved throughput without improving value.

Some harness behavior can move into the model

Harness-Zero, published September 21 and therefore used here as dated learning rather than fresh news, tests whether behavior induced by a specialized agent harness can survive when that harness is removed. A harnessing agent observes the optimized harness, corrects the student’s proposed actions into the target harness’s action space, and turns those corrections into fine-tuning trajectories. (Harness-Zero)

Across SpreadsheetBench, AppWorld, and a USPTO chemistry task, the distilled Qwen3.5-9B model averaged 44.3% success under a minimal harness. The base model scored 23.3% under the same harness. The base model with the evolved specialized harness scored 41.7%.

The macro-average is tempting, but the domain split is the design signal. The distilled model exceeded the evolved-harness setup on spreadsheets and AppWorld, while it trailed on the chemistry task. The paper reports 82.3% average recovery across 28 harness-induced behavior patterns, not complete transfer of every useful capability.

Procedural corrections can become training data. Domain tools, private knowledge, changing policy, credentials, context compaction, and effect controls still belong outside the weights. A practical release process should therefore ablate the boundary instead of arguing about it abstractly:

  1. Evaluate the base model with the minimal harness.
  2. Add the specialized harness and measure the marginal gain by task family.
  3. Distill only the behaviors that remain valid across policy and tool versions.
  4. Retest the distilled model with the minimal harness and with the production harness.
  5. Version the model and harness as a pair whenever their interaction changes.

Distillation can reduce routing complexity and token overhead. It also makes a behavior harder to inspect and update. The right boundary depends on its change rate. Stable procedural habits are candidates for weights. Revocable authority and customer-specific policy are not.

Prompt retrieval needs a failure contract per surface

Fin’s September 25 engineering account is also older than the news window, but it is useful application evidence for the same boundary. As customers add communication rules, data connectors, and multi-step procedures, Fin does not place the entire configuration into every prompt. It trains separate cross-encoder selectors from production-model decisions, then uses different retrieval contracts for each surface. (Fin)

Guidance uses a top-ten shortlist and fails open to the full set on an error. Connectors use a gate: if anything appears relevant, the full connector set reaches the planner because connectors may chain. Procedures use an absolute threshold because the correct answer is often no procedure at all. Fin reports that 61% of planning calls produced an empty plan and 84% of procedure checks found nothing applicable.

On a held-out week of production traffic, the reported selector recall was 0.97 for guidance and 0.95 for both connectors and procedures. The production model’s own full-context decisions served as the labels, so the evaluation measures fidelity to a teacher rather than independent correctness. Live tests kept resolution broadly stable for connector and procedure selectors. Guidance selection slightly reduced resolution while increasing adherence to customer rules. Fin says cost and latency fell but withholds exact effect sizes.

The useful pattern is not the undisclosed savings. It is that retrieval semantics follow the cost of a miss. A single top-k rule for every context surface would be simpler and less safe. Teams should classify prompt inputs before choosing a selector:

  • Advisory context can tolerate ranking and omission.
  • Policy context should fail open or route to a deterministic check.
  • Action capability needs high-recall gating and an authority check at execution.
  • Exclusive procedures need an explicit “none applies” outcome.

This turns context engineering into a set of contracts rather than a larger search box.

Stay Sharp: zero failures is not 100% reliability

Observed 20 out of 20 is a useful descriptive metric. It is not evidence that the true success probability is 100%. If independent attempts have success probability p, the probability of seeing twenty successes is p^20.

An exact one-sided 95% lower confidence bound after twenty successes and no failures is about 86.1%. In other words, twenty clean runs are still statistically compatible with a service that succeeds only about 86 times in 100. NIST documents exact binomial confidence limits for this setting and warns that normal approximations can be inaccurate with small samples or few failures. (NIST) (NIST handbook)

The sample grows quickly when the service objective is strict. With zero observed failures, demonstrating a success probability of at least 99% at one-sided 95% confidence requires 299 independent successes. That calculation assumes identically distributed, independent trials. Shared prompts, cached plans, reused state, or correlated provider incidents reduce the effective evidence.

This gives an evaluation owner four questions to answer before publishing “reliable”:

  1. What success probability must the service exceed?
  2. What confidence level supports the release decision?
  3. Which sources of variation are independently resampled?
  4. Which task families receive enough repetitions to detect their own failure rates?

The most useful shift this week is not from one leaderboard column to another. It is from asking whether an agent has ever succeeded to specifying how much repeat evidence a real decision requires. Task selection, model training, prompt retrieval, and terminal-state checks then become separate levers. Reliability improves when each lever has a contract and none is asked to compensate silently for the others.