In this article
OpenAI made Ultrafast mode generally available for GPT-6.1 Sol on October 8. The tier is selected per Responses API request with service_tier: "ultrafast", has rate limits separate from the standard and fast tiers, and supports US and EU data residency as well as global processing. OpenAI recommends a persistent WebSocket for tool-heavy agents so connection overhead does not consume the latency gain. (OpenAI changelog) (Ultrafast guide)
The price makes this an architecture decision rather than a harmless request flag. For a short-context GPT-6.1 Sol request, standard processing costs $2 per million input tokens and $10 per million output tokens. Ultrafast costs $12 and $60 respectively. A request with 10,000 uncached input tokens and 2,000 output tokens therefore moves from about $0.04 to $0.24 before tool fees or a regional-processing premium. (OpenAI pricing)
That premium can be rational when a delay changes task success, operator attention, or revenue. It is waste when token generation is not the critical path. The release should trigger a latency budget, not a fleet-wide tier change.
The tier accelerates one component of the task
OpenAI describes Ultrafast as reducing the time between generated output tokens. That is not the same as guaranteeing a lower end-to-end completion time. An agent request can wait in several places:
- admission and queueing before inference begins;
- time to the first useful output or tool call;
- generation between output tokens;
- network, database, browser, or human approval waits;
- retries, validation, and final review.
Measure those clocks separately. If generation occupies 1.5 seconds of an eight-second workflow, eliminating generation time entirely would save at most 18.75% in that illustrative run. A faster decoder cannot repair a slow search service, a serial tool plan, or an approval queue.
The denominator also matters. For interactive chat, time to the first useful sentence and inter-token smoothness may determine perceived quality. For an agent that makes twelve tool calls, the relevant measure may be time to a verified terminal state. For a coding workflow, reviewer minutes and rollback rate can dominate both.
Build the baseline from production-shaped traces, not a single prompt. Report at least p50 and p95 for first useful output, tool round trips, final completion, and failure recovery. Segment by prompt length, output length, tool count, region, and task family. Then the team can see whether the expensive tier changes the decision metric or only one attractive chart.
Route the tier at the response boundary
Because the service tier is chosen on each response, an orchestrator does not need one speed class for an entire job. It can reserve Ultrafast for steps where delay has a supported cost and keep other work on standard processing.
A practical policy might use standard processing for background extraction, batch enrichment, and speculative branches. It could promote a customer-visible synthesis, an incident-response decision, or a latency-sensitive tool turn only after the workflow reaches a state where faster generation changes the outcome. The exact split belongs to workload evidence, not these examples.
Four controls keep that routing honest:
- Quality parity: run the same acceptance set across tiers and connection modes. Faster delivery is not useful if timeout handling, truncation, or tool sequencing changes the result.
- Incremental economics: divide the additional tier cost by successful tasks whose latency objective improved, not by total requests sent.
- Capacity behavior: test separate rate limits, backpressure, and the application’s explicit fallback before a peak event.
- Regional evidence: verify latency and capacity from the required processing region. Residency eligibility does not establish a regional performance distribution.
This turns Ultrafast into a bounded resource. A hard budget can cap its use by task, tenant, or incident. An evaluation gate can prevent promotion when the measured completion-time gain is smaller than the premium warrants.
Model and tier routing need a decision record
Google’s October 8 Gemini agent announcement makes the routing problem broader. Google says the enterprise agent can choose among Gemini models and Anthropic’s Claude models, create temporary subagents with their own identities, and use smart routing and spend caps. It also says model choice is separate from the agent’s persistent context, skills, and data. These are provider descriptions of a new system, not independent evidence of reliability or savings. (Google Cloud)
The useful architecture question is what survives when a router changes the model or service class. For every consequential step, preserve:
- model family, exact snapshot, reasoning configuration, and service tier;
- the routing rule and evidence that selected them;
- prompt, tool, memory, and policy versions;
- residency, rate-limit, and fallback requirements;
- estimated and actual cost;
- task outcome, validation result, and human intervention.
Without that record, a cost or latency regression becomes difficult to localize. The router may have selected a different model, changed the tier, crossed a context threshold, lost a cache hit, or triggered a fallback. A trace that records only the visible agent name hides the decision that produced the bill and the behavior.
Do not optimize routing against average token latency alone. Give each task family a quality floor, a completion-time objective, and a maximum marginal cost. Let the router choose only among configurations that already pass the quality and policy gates. Speed is then an optimization inside an approved set, not permission to change the system silently.
Stay Sharp: Markdown is an intermediate representation, not ground truth
LlamaIndex published a useful document-processing note on October 8. It argues that Markdown is a practical intermediate format because headings, lists, links, and simple tables preserve more structure than plain text while remaining readable during debugging. It recommends HTML when merged table cells require rowspan or colspan, JSON for downstream structured records, and retained images or layout coordinates when text cannot preserve the needed spatial relationship. (LlamaIndex)
The mechanism matters for retrieval. A value extracted correctly can still become wrong evidence if its header, unit, footnote, or section boundary disappears. A chunker can use Markdown headings as boundaries and attach the section title to each chunk. A complex table may need normalized headers or an HTML representation so a year label remains attached to the right revenue and margin columns.
Treat the representation as a testable contract:
- choose documents with known merged headers, footnotes, lists, and diagrams;
- assert that every value retains its labels, units, and provenance;
- compare the intermediate representation with the source page, not only with expected answer text;
- keep page and bounding-box metadata for review and citation;
- measure retrieval and extraction failures by structural cause.
Markdown makes the pipeline inspectable. It does not prove that the parser reconstructed the document correctly. That distinction is the same one a routing system needs: a convenient interface must preserve enough underlying evidence to explain a decision.
What to watch
- Whether OpenAI publishes task-relevant latency distributions for Ultrafast, rather than only tier availability and token prices.
- How separate rate limits and explicit application fallbacks behave under peak traffic.
- Whether enterprise routers expose the chosen model, tier, reason, and policy decision in exportable traces.
- Whether document pipelines evaluate structural fidelity before retrieval and generation scores hide an upstream parsing error.