In this article
Anthropic changes the default execution model for new Cowork tasks on Claude Pro and Max today. Those tasks now run in Anthropic’s cloud, and the Only on your computer setting is removed for those plans. Existing local tasks can continue where they started. The stated benefits are continuity across desktop, web, and mobile, background work after a laptop closes, and scheduled tasks with no device online. (Anthropic Help Center)
This is not a routine hosting change. It moves the agent loop and code execution across a trust boundary while preserving routes back to files, browsers, apps, and connectors. Teams that evaluated Cowork as a local execution product should reopen that evaluation. The correct question is no longer simply, “Is the agent sandboxed?” It is, “Which capability crosses which boundary, under whose policy, and with what evidence?”
The October 6 change is specific to new Pro and Max tasks. Team and Enterprise controls differ, so administrators should verify their own plan and configuration rather than generalize the consumer-plan migration.
The sandbox moved, but the system did not become purely remote
Anthropic says each cloud session receives an isolated temporary sandbox on its infrastructure. Sandboxes do not share state, cannot reach private, link-local, cloud-metadata, or Anthropic-internal addresses, and send permitted traffic through an external egress proxy. Connector authorization tokens remain outside the sandbox, and session credentials expire within hours. (Cowork architecture overview)
Those are meaningful controls, but they cover one execution environment. A cloud session can still request a local file or browser through the Claude Desktop app. Local file access is limited to connected folders and checked against the user’s permissions. If the desktop app is offline, the cloud session cannot reach the device. Files opened through that bridge are processed on Anthropic’s servers. (Cowork architecture overview)
Other paths have different properties. Anthropic’s safety guidance says connector and file operations use permission checks, while computer use acts through the screen and has no sandbox between the agent and what is visible there. A task can also move data between connected applications. The company recommends human oversight when actions are hard to reverse and warns against sensitive work in browser-driven flows. (Use Cowork safely)
The architectural mistake is to draw one box around “Cowork” and assign that box a single risk rating. The product is a graph of execution and capability paths:
- Cloud code: isolated on Anthropic infrastructure, with server-enforced network policy.
- Local bridge: selected folders and desktop tools, available only while the desktop app is connected.
- Server-side connectors: authorized services called outside the task sandbox.
- Computer use: screen-level interaction with a distinct approval and observability model.
- Scheduled work: cloud tasks that can start or continue without an attended device.
Each path needs its own data classification, action policy, logging expectation, and emergency stop. A sandbox verdict cannot substitute for that inventory.
Make execution locality part of the task contract
Before enabling recurring or higher-impact work, record five fields for every task class:
- Execution location: cloud session, local session where still supported, or a mixed path through the desktop bridge.
- Reachable capabilities: folders, connectors, network destinations, browser surfaces, plugins, and computer use.
- Authority: the user or service identity, allowed actions, approval requirements, and revocation owner.
- Data movement: where inputs are stored, where processing occurs, what crosses the device boundary, and how long artifacts remain.
- Evidence: session identifiers, tool-call records, output provenance, policy decisions, and the logs available for incident review.
Treat a change to any field as a new deployment revision. A workflow approved for a local VM should not silently inherit approval for a cloud sandbox plus desktop bridge. Conversely, cloud execution can improve continuity and central policy enforcement, but only if the evidence and control plane cover the whole route.
Anthropic’s earlier containment account explains why this distinction matters. Its first Cowork design used a sealed local VM so the agent could see only mounted workspaces. The company later described two limits: isolation can hide activity from host endpoint monitoring, and an approved network destination can still expose a dangerous capability. In one disclosed case, access to an allowed Anthropic API domain enabled file exfiltration with an attacker-controlled key until Anthropic added token-aware proxy checks. (Anthropic Engineering)
The lesson is not that local or cloud execution is categorically safer. It is that enforcement belongs at the narrowest capability boundary. Domain allowlists need request-level semantics. Folder permission needs resolved-path checks. Screen interaction needs a separate policy from connector calls. Central execution needs logs that compensate for reduced endpoint visibility.
Application watch: define the exit from embedded engineering
Atlassian published a Forward Deployed Engineering playbook on October 4 based on more than 100 enterprise engagements. It reports over 80 production agents and roughly 12 weeks to self-sufficient customer teams. The operating model starts with a measurable business problem, maps actors, systems, handoffs, and permissions, then co-builds the workflow with the customer team. (Atlassian playbook)
The useful design point is the exit condition. Atlassian describes four phases: discover, build, adopt, and value. It says engagements leave behind working solutions and reusable patterns rather than a permanent expert dependency. That suggests a practical readiness gate for any embedded AI team: the receiving team can operate, evaluate, revoke, and change the system without the original builders. (Atlassian FDE)
The outcome claims need caution. Atlassian cites large savings and cycle-time reductions, but the public pages identify customers only by sector and do not publish measurement protocols, baselines, or counterfactuals. Use the workflow pattern as an implementation hypothesis, not the figures as a benchmark. Require named metric owners, pre-deployment baselines, failure costs, and a date when the customer team must demonstrate independent operation.
Stay Sharp: skill selection is a retrieval system
SkillSeek, published September 30, tests a two-stage retriever for agent skills: a BGE bi-encoder proposes candidates, then a small cross-encoder reranks them. On 89 SkillsBench tasks across four pool and backbone settings, the authors report that BM25 matched or exceeded an LLM-mediated retrieval loop in three settings, while the cross-encoder closed the remaining gap. Their reported per-trial spend fell from $51.30 to $27.54, close to the no-skill baseline. (Yang et al.)
The mechanism is ordinary information retrieval. First-stage recall sets a ceiling. Reranking can improve ordering only among skills that the first stage retrieved. That gives teams three distinct measurements:
- recall of the skill that an expert would want in the candidate set;
- ranking quality among the returned candidates;
- downstream task success after the agent uses the selected skill.
Do not collapse them into one pass rate. If recall is low, improve skill descriptions, indexing, and query formulation. If ranking is weak, tune the reranker. If retrieval is sound but tasks still fail, inspect the skill’s instructions and the execution harness.
The evidence is bounded. The result comes from SkillsBench, OpenHands, and the tested pools and backbones. It does not establish that BM25 is sufficient for proprietary catalogs with sparse descriptions, access controls, or organization-specific vocabulary. The transferable lesson is narrower: establish a cheap lexical baseline and measure the retrieval ceiling before putting an LLM loop on every task.
What to do this week
- Inventory Cowork task classes by execution location and capability path, with Pro and Max separated from Team and Enterprise.
- Reapprove workflows whose processing location or unattended behavior changed, especially those that touch regulated or confidential files.
- Test revocation and incident reconstruction for the cloud sandbox, desktop bridge, connectors, and computer use as separate paths.
- Require an operational handoff test before an embedded AI team declares a workflow self-sufficient.
- Benchmark skill retrieval in three layers: candidate recall, reranking quality, and downstream task success.