In this article 5 sections

Ironclad and OpenAI have turned contracting work into 11 agent research tasks spanning legal, commercial, and procurement workflows. Each task has 8 to 50 scoring criteria and runs in a hosted Ironclad environment. GPT-6 Astra averaged 55.0% across those criteria, compared with 41.6% for GPT-5.6 Sol. The estimated time per attempt fell from 37.0 minutes to 19.2 minutes. (OpenAI and Ironclad)

The percentages are not the most useful result. Eleven tasks cannot represent every contracting workflow, and the reported time is a simulation rather than observed customer labor. The public account does not provide repeat-run variance, failure costs, or evidence that the model improved a production outcome.

The more transferable contribution is the evaluation artifact. Domain experts converted business rules into tasks, terminal states, and explicit criteria. That is the unit an engineering team can version, test, and review before granting an agent real authority.

Turn policy prose into executable acceptance criteria

Contract workflows fail in ways that a generic “task completed” score will miss. An agent can upload the right agreement under the wrong counterparty, select an acceptable clause but skip the required approval, or reach a completed screen while leaving a material field inconsistent. A useful task specification has to make those distinctions visible.

For each workflow, define four layers:

  1. Initial state: the contract, counterparty, policy version, permissions, and application state the agent receives.
  2. Allowed transitions: the actions it may take, the approvals it must obtain, and the conditions that make an action legal.
  3. Terminal state: the exact records, fields, attachments, and workflow status that must exist when the task ends.
  4. Evidence: the application events, screenshots, tool calls, and policy decisions needed to prove each criterion.

Ironclad and OpenAI used synthetic tasks based on public SEC EDGAR contracts, filtered to remove personal information. That is a sensible starting point for reproducibility, but it also bounds the result. Public filings do not reproduce a company’s private playbook, messy intake data, negotiated exceptions, or historical process debt. A production corpus needs representative internal cases, including rare exceptions, with appropriate access and retention controls.

The score should also preserve the shape of failure. A missed optional field and an unauthorized approval bypass cannot both become one lost point. Record at least:

  • criterion pass or fail;
  • policy severity and reversibility;
  • omitted, incorrect, or extra action;
  • human intervention required;
  • evidence quality;
  • result stability across repeated attempts.

An aggregate score remains useful for trend detection. It should never be the only release gate.

Evaluate the workflow, not the mouse path

Computer-use agents introduce a temptation to score visible interactions. Clicking the expected controls is not the business outcome. The stable contract is usually the resulting system state plus the evidence that required policy transitions occurred.

Separate three tests:

  • Policy test: Did the proposed action comply with the current business rule?
  • Execution test: Did the agent produce the required application state without prohibited side effects?
  • Review test: Can an operator reconstruct why the action was allowed and what changed?

This separation makes the evaluation portable across models and interface revisions. If Ironclad changes a menu or a model uses a different navigation path, the business criteria can remain stable. If the approval threshold changes, the policy version changes even when the interface does not.

Before a release, run a fixed regression set and a rotating exception set. The fixed set detects model or prompt regressions. The rotating set prevents teams from optimizing only for familiar fixtures. For high-impact actions, require repeated success and zero critical-policy failures rather than accepting a high mean that hides one catastrophic path.

The research collaboration reports that an internal model reached 63.7%, above the public GPT-6 Astra result. That is evidence of headroom inside this task design, not evidence that the workflow is production-ready. The remaining criteria still need classification by severity and cause.

Model radar: Mistral Large 4 is a preview, not yet a portable artifact

Mistral released the Mistral Large 4 API in public preview on October 6. Its launch post describes a one-trillion-parameter mixture-of-experts model with 49 billion active parameters, while the linked model page showed 52 billion active parameters during editorial review. Both give a one-million-token context window. The company says model weights will arrive at the end of October. (Mistral, Mistral Docs)

That timing matters for architecture decisions. Teams can evaluate the hosted API now, but they cannot yet verify the promised open-weight deployment path, inspect the released artifact, or test self-hosted operational characteristics. Treat “weights later” as a pending release condition rather than present portability.

The same task corpus should follow the model. Run Mistral Large 4, GPT-6 Astra, and future candidates against identical policy versions, evidence requirements, and failure weights. Vendor benchmarks can inform discovery. They should not replace the workflow-specific gate.

Stay Sharp: generate sequences, then assert invariants

Hypothesis’s rule-based state machines offer a useful mental model for testing business workflows. Instead of supplying one input to one test, you define primitive actions that Hypothesis can combine into sequences. Bundles pass generated values between steps, preconditions restrict which actions are valid in the current state, and invariants run after every step. When a sequence fails, Hypothesis tries to shrink it into a smaller reproducer. (Hypothesis documentation)

Map that mechanism onto contract operations:

  • rules create an intake request, change a deal value, add a clause, request approval, or submit an agreement;
  • preconditions prevent submission before required fields exist;
  • bundles reuse the same counterparty, agreement, and approver across later actions;
  • invariants assert that no above-threshold agreement becomes final without the required legal and finance approvals.

This catches failures that isolated happy-path tasks miss, especially when an earlier action changes the meaning of a later one. It also produces a short sequence that an engineer and domain owner can inspect together.

The limit is the model you write. Generated sequences explore only the actions, state, and invariants represented in the harness. They do not prove that a legal judgment is correct, that a user interface was interpreted faithfully, or that an omitted business rule is safe. Domain experts still own the policy model and the severity of each invariant.

What to do this week

  • Choose one consequential workflow and express its initial state, legal transitions, terminal state, and evidence as a versioned task.
  • Split aggregate performance into criterion severity, failure type, intervention cost, and repeat stability.
  • Gate deployment on critical-policy invariants and representative exceptions, not a mean score alone.
  • Keep workflow tests stable across model candidates so vendor claims are compared against the same business contract.
  • Evaluate the Mistral Large 4 API if relevant, but defer portability conclusions until the promised weights, license, and deployment artifacts are available.