In this article 6 sections

Siemens published one of the more useful enterprise-agent deployment accounts this week. Salesforce says two agents now handle engagement and qualification for roughly 2,500 in-scope inbound leads per month across 132 countries. The qualification agent talks naturally with prospects, but the final routing decision is governed by about 50 predefined rules. A lead with at least two positive BANT responses is sent to a seller; lower-scoring leads follow a predefined closure path. (Salesforce)

That boundary is more interesting than the chatbot itself. The model handles language and information gathering. CRM fields remain authoritative. Deterministic policy decides what happens next.

Two operational updates deserve attention beside it. OpenAI reported that customers using the Agents API experienced delays or could not start turns in managed sessions on September 14 before service recovered later that day. And DeepSeek has reversed its planned September 14 retirement of V4 Pro, leaving the model available with unchanged billing. (OpenAI Status) (DeepSeek)

Siemens shows a practical form of bounded autonomy

The Siemens workflow starts with an engagement agent that contacts inbound leads and can follow up when they do not respond. Prospects who engage move to a qualification agent that collects budget, purchasing authority, need, and timeline information and writes the answers into Sales Cloud as the conversation progresses. Agent Script then applies Siemens’ qualification policy. (Salesforce)

Salesforce reports that the system now covers 100% of the approximately 2,500 in-scope monthly leads, has cut response time from days to minutes, and has produced an 11% engagement rate and 6% qualification rate. It also reports that 80% of interactions receive a Good or Very Good rating. The customer story does not provide the measurement window, denominators, rating instrument, or comparison design. These are vendor-hosted customer results, not an independent controlled evaluation, so they should not be generalized as benchmark evidence. (Salesforce)

The architecture is still worth copying conceptually.

A production agent does not need probabilistic control over every stage of a workflow. It can use the model where uncertainty is genuinely linguistic or contextual, then hand decisions to explicit policy when the business rule is stable enough to encode. In this case:

  1. the model conducts the conversation;
  2. structured CRM fields record the relevant facts;
  3. deterministic rules evaluate those facts;
  4. the system routes or closes the lead;
  5. a seller receives the context only when the policy says human attention is warranted.

This split matters for evaluation. You can evaluate conversational quality separately from qualification correctness. You can inspect whether BANT fields were extracted correctly, whether rules executed as specified, whether the right leads reached sellers, and whether exceptions clustered around a particular policy boundary.

The rule layer does not make the end-to-end qualification result deterministic. A model or prompt change can still alter which BANT fields are populated or how ambiguous answers are normalized, and therefore change routing even when the rules stay fixed. Version the extraction layer, retain the evidence behind each field, and test field accuracy and routing outcomes together.

The separation still creates a safer change process. The qualification threshold can change without retraining the conversational layer, while model changes can be evaluated without rewriting the business policy. The consequential decision becomes more inspectable, but only if teams test the probabilistic inputs as well as the deterministic rules.

The Agents API outage turns managed orchestration into an SLO question

OpenAI’s status page says that from 1:30 PM Pacific Time on September 14, customers using the Agents API experienced delays or were unable to start turns in managed sessions. OpenAI applied mitigations and later reported that managed sessions were processing normally again. (OpenAI Status)

This is a material update to last week’s Agents API launch because it demonstrates a new failure boundary in production rather than another capability claim.

A managed agent runtime can absorb session persistence, context compaction, tool discovery, orchestration, and recovery logic. That removes application code, but it also means the ability to start or advance a turn becomes a provider dependency with its own SLO.

For an application team, ordinary API uptime is not enough. Useful service indicators include:

  • successful turn-start rate;
  • time spent waiting for a managed session to accept work;
  • age of queued work;
  • percentage of sessions that recover without user-visible duplication;
  • and the fraction of business tasks that can degrade gracefully when orchestration is unavailable.

The durable business intent should also exist outside the hosted agent session. If a user asked to review an account, prepare a report, or execute a multi-step workflow, your system should still know that the work exists when a provider turn cannot start. That does not imply reimplementing the provider’s agent runtime. It means keeping enough authoritative task state to explain, retry, cancel, or reroute work deliberately.

Managed orchestration is useful precisely because it centralizes hard runtime problems. Treat it like any other stateful platform dependency: define the failure semantics before the first incident.

DeepSeek reversed a model retirement before the deadline

DeepSeek originally announced that requests to deepseek-v4-pro would be routed to V4.1 Flash after 12:00 Beijing Time on September 14. Its current changelog now says that, in response to user demand, V4 Pro will continue after September 14 with unchanged billing, and that further notice will be provided if this changes. (DeepSeek)

This is a small but useful model-lifecycle lesson. A deprecation calendar is not immutable configuration. Providers can accelerate a retirement, delay it, replace an endpoint silently, or reverse course entirely.

During this research pass, some indexed DeepSeek documentation still reflected the earlier routing plan while the current changelog reflected the reversal. That is exactly the kind of transition where a spreadsheet of remembered dates is weaker than a monitored lifecycle record.

For production dependencies, track both the announced future state and the currently observed contract. Before a scheduled migration boundary, re-read the primary source, run a small behavioral check against the exact model name, and retain enough evaluation evidence to detect a silent substitution even when the endpoint keeps returning HTTP 200.

Stay Sharp: multi-stage retrieval is a fidelity-allocation problem

LlamaIndex published a useful engineering note on September 11 describing a two-pass document workflow for ad hoc data rooms. It is older than today’s formal news window, so treat it as learning material rather than fresh news. (LlamaIndex)

The pattern is simple:

  1. run a cheap first pass over all documents to obtain searchable text;
  2. retrieve the small set of pages likely to matter;
  3. apply expensive VLM-based OCR only to those pages.

In one FinanceBench example, LlamaIndex processed 12,013 pages across 84 filings with a lightweight first pass, then used higher-fidelity OCR on only two pages needed for the question. The numbers are vendor-reported and the post also promotes LlamaIndex products, but the architecture is broadly useful. (LlamaIndex)

The important abstraction is not OCR. It is selective fidelity.

Many AI systems have stages with very different cost and accuracy profiles: lexical search versus dense retrieval, small reranker versus frontier model, low-resolution vision versus detailed perception, heuristic filter versus expensive verifier. The system should spend the expensive stage where it changes the decision.

That produces two distinct evaluation questions:

  • Recall of the cheap stage: did it preserve enough candidates for the expensive stage to recover the right answer?
  • Value of the expensive stage: when invoked, did the extra fidelity actually improve correctness enough to justify its cost and latency?

A cascade fails if the first pass drops evidence that no later stage can recover. This is why LlamaIndex recommends the two-pass pattern for roughly 10 to 100 ad hoc documents, while favoring higher-quality preprocessing up front for large reusable corpora where retrieval quality depends on the stored representation. (LlamaIndex)

The senior-level mental model is to treat retrieval as a budgeted evidence pipeline, not a single top-k call. Allocate fidelity where uncertainty and consequence justify it.

Watchlist

  • Sora API: OpenAI still marks the API for permanent shutdown on September 24, 2026. Migration and data-export work should already be in execution, not planning. (OpenAI deprecations)
  • Agents API reliability: watch whether OpenAI publishes more detailed incident scope or managed-session reliability guidance after the September 14 degradation.
  • DeepSeek V4 Pro: the retirement was reversed, not permanently ruled out. Keep the dependency on the lifecycle watchlist.
  • Enterprise agent evidence: look for Siemens publishing longer-term seller outcomes, false-positive qualification rates, or exception handling outside the Salesforce customer-story format.

Further reading