In this article
OpenAI launched Dots and the Decisions API on September 29. The important architecture signal is not simply that one agent can now work for longer. The two products expose three materially different authority shapes: background discovery through read-only tools, actions governed by approval rules, and decisions constrained to a finite answer set. (OpenAI Dots) (OpenAI DevDay recap)
That separation is useful even for teams that never adopt either product. A system that observes a queue, recommends a refund, and executes the refund should not inherit one undifferentiated permission called “agent access.” Observation, judgment, and side effects have different failure costs, evidence needs, and rollback options.
Dots makes background work a distinct permission class
OpenAI describes Dots as always-on agents powered by GPT-6 Astra, with a cloud computer and connections to more than 4,000 apps. A Dot can carry context across ChatGPT, Slack, and Teams, work on several projects, and send progress updates or questions. The product is rolling out to Pro and Business Premium users in eligible markets, while Enterprise, Edu, and Healthcare access is an admin-enabled beta. (OpenAI Dots)
The more consequential detail is how the product divides passive work from action. When a user is not actively working with it, a Dot can perform what OpenAI calls proactive research. In that mode, connected-app tools are restricted to read-only use. They cannot send messages, modify app content, or control a browser or computer. A separate action path applies built-in rules and user-defined Custom Rules that allow, require approval for, or block actions. Auto-review checks proposed account or information-sharing actions against those rules and safety requirements. Password changes remain with the user. (OpenAI Dots)
Those are vendor-described controls, not an independent measurement of their effectiveness. Reuters independently confirmed the launch, the Slack and Teams integrations, and OpenAI’s claim that sensitive tasks need explicit consent. It also placed the release against recent agent incidents and OpenAI’s decision a day earlier to withhold a more capable Astra variant over safety concerns. (Reuters) AP reported that OpenAI is investing more in agent safety, security, and monitoring, but the public launch material does not provide false-approval rates, policy-bypass rates, or recovery results for Dots. (AP)
For an enterprise deployment, preserve the distinction in the local control plane:
- Observe: grant narrowly scoped, read-only access with explicit data boundaries and a retention policy.
- Propose: require the agent to emit the intended effect, affected object, evidence, and policy revision before approval.
- Act: issue a short-lived capability for one effect, not a reusable credential for an entire app.
- Verify: record the actual side-effect receipt separately from the model’s statement that the action succeeded.
A progress feed is not an audit log. The durable record should join the goal, source data, rule version, requested effect, approval identity, execution receipt, and any compensation step. Otherwise, teams can see that an agent was busy without being able to reconstruct why a consequential change occurred.
Finite decisions deserve a narrower runtime
The Decisions API takes a different route. OpenAI says developers define specific questions and finite possible answers, supply text or image context, and receive a selection that can classify content, route a request, or choose an agent’s next action. The API is in limited preview. (OpenAI DevDay recap)
That is not merely a smaller user interface around a general agent. It turns the output space into part of the contract. A fixed answer set can be schema-validated, covered with per-class tests, mapped to deterministic downstream handlers, and denied when evidence is insufficient. It also lets operators measure abstention, confusion, and cost by decision class rather than reviewing free-form text after the fact.
The constraint does not make a decision correct. Teams still need calibrated thresholds, an explicit unknown or escalate state, and tests built from actual edge cases. The downstream handler must also enforce authorization independently. A model selecting approve is evidence for a policy decision, not the permission that executes it.
The practical design choice is therefore about authority geometry:
- Use a read-only observer when the value comes from detecting change.
- Use a finite decision service when the business process already has a closed set of outcomes.
- Use a permissioned agent when the task genuinely requires open-ended planning and tool use.
- Split one workflow across those modes instead of granting the broadest mode to every step.
This reduces the amount of behavior that must be evaluated as open-ended. It also makes escalation a designed transition between modes rather than a prompt asking the same agent to be careful.
A deadline today: Gemini Omni Flash preview
Google’s Gemini API changelog says gemini-omni-flash-preview reaches its deprecation date on September 30. The listed GA replacement is gemini-omni-1.1-flash, released on August 27 with video extension, first-to-last-frame interpolation, and resolution controls. (Gemini API changelog)
Before moving remaining traffic, compare request schemas, output artifacts, safety behavior, latency, and cost on production-shaped fixtures. Treat the endpoint change as a model and media-pipeline migration, not a string replacement. Pin the target model name in configuration, log it with every generated asset, and keep a rollback decision that does not depend on the deprecated endpoint remaining available.
Stay Sharp: provenance is a relationship, not a citation string
Multiverse Computing researchers published a practical lesson for multi-tool agents on September 29. Their ProvenanceGuard work targets cross-source conflation: a statement is supported somewhere in the combined evidence, but the answer attributes it to the wrong source. A customer policy might establish a 30-day refund window while the agent claims that the customer’s account record established it. A source-blind faithfulness check can pass because the fact exists in the pooled context. (ProvenanceGuard)
The proposed verifier keeps each tool output’s source identity through claim decomposition, source routing, support checking, attribution checking, and an allow-or-block decision. In the authors’ held-out medical test, experts marked 139 of 361 claims as claims that should not pass. The system caught 138, but it also held 67 claims that experts considered supported. It identified the right source about 86 percent of the time when a source was identifiable. In a harder test with several similar sources, exact source identification fell to 50.3 percent. (ProvenanceGuard)
The repair results expose the operational trade-off. The full-trace run resolved 173 blocked answers, but 144 ended in fallback text rather than a substantive repair. That is a defensible bias for sensitive work, yet it also means the method can exchange silent attribution errors for visible non-answers. The study used medical traces and a conservative local-model configuration, so teams should recalibrate it for their own source types and review costs.
The reusable mechanism is simpler than the full implementation: keep a stable source ID on every tool result, require every material claim to point to the source that supports it, and test attribution separately from factual support. For an always-on agent, that relationship should survive from observation through decision and into the side-effect receipt. A citation added at the end cannot reconstruct provenance that the workflow discarded earlier.
What to watch
- Whether OpenAI publishes measured approval, intervention, and recovery behavior for Dots rather than only control descriptions.
- Whether specialist Dots expose independent identities, short-lived credentials, rule revisions, and exportable action receipts to enterprise administrators.
- Whether the Decisions API includes native abstention, calibration evidence, and immutable decision schemas as it moves beyond limited preview.
- Whether Google clarifies shutdown behavior for calls that still target
gemini-omni-flash-previewafter September 30. - Whether source-aware verification can distinguish highly similar records without the conservative blocking rate becoming operationally prohibitive.