<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Critical Path AI</title>
    <link>https://criticalpathai.dev/</link>
    <description>Daily and weekly analysis of AI systems, models, agents, research and engineering, ranked by consequence rather than hype.</description>
    <language>en</language>
    <lastBuildDate>Thu, 08 Oct 2026 00:00:00 GMT</lastBuildDate>
    <atom:link href="https://criticalpathai.dev/rss.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Fin&apos;s 8.6% hard-resolution lift needs a feedback denominator</title>
      <link>https://criticalpathai.dev/daily/2026-10-08/</link>
      <guid isPermaLink="true">https://criticalpathai.dev/daily/2026-10-08/</guid>
      <pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate>
      <description>Fin changed its base model and post-training system while reporting a lift in explicit user-confirmed resolutions. The transferable lesson is how to separate system improvement from model attribution and correct for the users who never answer the survey.</description>
      <category>Daily Pulse</category>
      <category>Production Evaluation</category>
      <category>Customer Support AI</category>
      <category>Feedback Systems</category>
      <category>Model Post-Training</category>
      <category>Causal Attribution</category>
    </item>
    <item>
      <title>Ironclad turned 11 workflows into agent tests. The rubric is the release artifact</title>
      <link>https://criticalpathai.dev/daily/2026-10-07/</link>
      <guid isPermaLink="true">https://criticalpathai.dev/daily/2026-10-07/</guid>
      <pubDate>Wed, 07 Oct 2026 00:00:00 GMT</pubDate>
      <description>OpenAI and Ironclad converted contracting work into 11 research tasks with explicit acceptance criteria. The transferable lesson is how to make business rules executable, measurable, and reviewable before an agent reaches production.</description>
      <category>Daily Pulse</category>
      <category>Agent Evaluation</category>
      <category>Business Workflow Automation</category>
      <category>Computer Use</category>
      <category>Contract Operations</category>
      <category>Model Testing</category>
      <category>Stateful Testing</category>
    </item>
    <item>
      <title>Claude moved Cowork&apos;s execution boundary. Rebuild the trust map</title>
      <link>https://criticalpathai.dev/daily/2026-10-06/</link>
      <guid isPermaLink="true">https://criticalpathai.dev/daily/2026-10-06/</guid>
      <pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate>
      <description>Cowork&apos;s Pro and Max cloud migration changes where code runs, where files are processed, and which controls remain on the device. Treat execution locality as governed configuration, not product trivia.</description>
      <category>Daily Pulse</category>
      <category>Agent Security</category>
      <category>Cloud Sandboxes</category>
      <category>Data Governance</category>
      <category>Desktop Integration</category>
      <category>Agent Operations</category>
      <category>Information Retrieval</category>
    </item>
    <item>
      <title>GPT-Rosalind starts billing today. Price the whole scientific run</title>
      <link>https://criticalpathai.dev/daily/2026-10-05/</link>
      <guid isPermaLink="true">https://criticalpathai.dev/daily/2026-10-05/</guid>
      <pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate>
      <description>GPT-Rosalind&apos;s metered research access and Paper2Agent&apos;s conversion failures show why scientific AI should be bought, governed, and audited as a reproducible run bundle.</description>
      <category>Daily Pulse</category>
      <category>Scientific AI</category>
      <category>Research Reproducibility</category>
      <category>Model Economics</category>
      <category>AI Governance</category>
      <category>MCP</category>
    </item>
    <item>
      <title>Kimi-K3 solved 94% of workflows once. Only 13% survived every run</title>
      <link>https://criticalpathai.dev/weekly/2026-w40/</link>
      <guid isPermaLink="true">https://criticalpathai.dev/weekly/2026-w40/</guid>
      <pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate>
      <description>New evidence separates occasional agent capability from dependable operation, while applied teams show where task choice, model training, and runtime retrieval each belong.</description>
      <category>Weekly Review</category>
      <category>Agent Evaluation</category>
      <category>Reliability Engineering</category>
      <category>AI for Science</category>
      <category>Agent Harnesses</category>
      <category>Context Engineering</category>
    </item>
    <item>
      <title>Gemini 4 Argon&apos;s coding score needs a benchmark audit</title>
      <link>https://criticalpathai.dev/daily/2026-10-02/</link>
      <guid isPermaLink="true">https://criticalpathai.dev/daily/2026-10-02/</guid>
      <pubDate>Fri, 02 Oct 2026 00:00:00 GMT</pubDate>
      <description>Google&apos;s restricted Gemini 4 Argon release shows why coding-model selection needs task-level verifier checks, production-shaped acceptance tests, and measured review cost.</description>
      <category>Daily Pulse</category>
      <category>Model Evaluation</category>
      <category>AI Engineering</category>
      <category>Benchmarking</category>
      <category>Software Engineering</category>
      <category>Model Lifecycle</category>
      <category>Agent Security</category>
      <category>Distillation</category>
    </item>
    <item>
      <title>OpenAI gave its always-on agent a read-only research mode</title>
      <link>https://criticalpathai.dev/daily/2026-09-30/</link>
      <guid isPermaLink="true">https://criticalpathai.dev/daily/2026-09-30/</guid>
      <pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate>
      <description>OpenAI&apos;s DevDay releases turn authority design into a product decision: continuous discovery, constrained judgment, and consequential action should not share one permission boundary.</description>
      <category>Daily Pulse</category>
      <category>Agent Architecture</category>
      <category>Autonomous Agents</category>
      <category>Decision Systems</category>
      <category>AI Governance</category>
      <category>MCP</category>
      <category>Model Deprecations</category>
    </item>
    <item>
      <title>NVIDIA put the agent watchdog outside the host</title>
      <link>https://criticalpathai.dev/daily/2026-09-29/</link>
      <guid isPermaLink="true">https://criticalpathai.dev/daily/2026-09-29/</guid>
      <pubDate>Tue, 29 Sep 2026 00:00:00 GMT</pubDate>
      <description>OpenShell 0.1.0 separates an agent workload from the supervisor that holds policy and credentials, while NVIDIA&apos;s optional Sentry design moves a second enforcement layer onto a BlueField DPU. The architecture improves failure isolation, but it does not solve policy correctness or prove containment effectiveness.</description>
      <category>Daily Pulse</category>
      <category>Agent Security</category>
      <category>Runtime Isolation</category>
      <category>Policy Enforcement</category>
      <category>AI Architecture</category>
      <category>Model Migration</category>
    </item>
    <item>
      <title>OpenAI&apos;s four-model shutdown hides three migrations behind one replacement</title>
      <link>https://criticalpathai.dev/daily/2026-09-28/</link>
      <guid isPermaLink="true">https://criticalpathai.dev/daily/2026-09-28/</guid>
      <pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate>
      <description>Four legacy model IDs point to GPT-5.6 Terra, but their callers do not share an interface, behavior contract, or cost baseline. Treating the change as a string replacement can hide the real cutover risk.</description>
      <category>Daily Pulse</category>
      <category>Model Lifecycle</category>
      <category>API Migration</category>
      <category>Evaluation</category>
      <category>AI Economics</category>
      <category>Reliability Engineering</category>
    </item>
    <item>
      <title>Meta’s agent handed calls to people. Delegation needs its own control plane</title>
      <link>https://criticalpathai.dev/weekly/2026-w39/</link>
      <guid isPermaLink="true">https://criticalpathai.dev/weekly/2026-w39/</guid>
      <pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate>
      <description>Muse exposed a specific agent boundary: authenticating the visible service does not reveal who performed an action, what authority crossed the handoff, or who owes the user a remedy.</description>
      <category>Weekly Review</category>
      <category>Agent Architecture</category>
      <category>Identity and Authorization</category>
      <category>AI Governance</category>
      <category>Agentic Commerce</category>
      <category>Human Oversight</category>
    </item>
    <item>
      <title>Anthropic now bills some refusals with zero output tokens</title>
      <link>https://criticalpathai.dev/daily/2026-09-25/</link>
      <guid isPermaLink="true">https://criticalpathai.dev/daily/2026-09-25/</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 GMT</pubDate>
      <description>A refusal can now return HTTP 200, incur a charge, and leave the task unfinished. LangChain&apos;s identity-scoped memory and Google&apos;s local-agent route show why production systems need separate outcome, authority, and execution ledgers.</description>
      <category>Daily Pulse</category>
      <category>API Economics</category>
      <category>Failure Taxonomy</category>
      <category>Agent Identity</category>
      <category>Local Inference</category>
      <category>AI Architecture</category>
    </item>
    <item>
      <title>Google says user devices will hold the keys to persistent AI memory</title>
      <link>https://criticalpathai.dev/daily/2026-09-24/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-24/</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 GMT</pubDate>
      <description>Google’s proposed Private AI Compute memory separates encrypted cloud storage, device-held keys, and attested enclaves, making recovery and deletion part of the product contract.</description>
      <category>Daily Pulse</category>
      <category>AI Memory</category>
      <category>Privacy Engineering</category>
      <category>Trusted Execution</category>
      <category>Key Management</category>
      <category>AI Architecture</category>
    </item>
    <item>
      <title>Changing an agent’s tools now has a cache bill and a replay contract</title>
      <link>https://criticalpathai.dev/daily/2026-09-23/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-23/</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 GMT</pubDate>
      <description>OpenAI and Anthropic changed how long-running agents preserve cached context, tools, and reasoning, turning configuration edits and model swaps into state migrations.</description>
      <category>Daily Pulse</category>
      <category>AI Agents</category>
      <category>Prompt Caching</category>
      <category>Model Migration</category>
      <category>Reliability</category>
      <category>AI Economics</category>
    </item>
    <item>
      <title>A typed evaluator matched five labels across 500 repeated decisions</title>
      <link>https://criticalpathai.dev/daily/2026-09-21/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-21/</guid>
      <pubDate>Mon, 21 Sep 2026 00:00:00 GMT</pubDate>
      <description>A narrow LangChain test suggests a useful evaluator layer between deterministic checks and generative judges, but low variance is not the same as trustworthy judgment.</description>
      <category>Daily Pulse</category>
      <category>Evaluation</category>
      <category>Agents</category>
      <category>AI Infrastructure</category>
      <category>MLOps</category>
    </item>
    <item>
      <title>AI teams are starting to measure work, not usage</title>
      <link>https://criticalpathai.dev/weekly/2026-w38/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/weekly/2026-w38/</guid>
      <pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate>
      <description>OpenAI and Anthropic exposed richer task-level metrics this week. The architectural opportunity is a task ledger that connects model activity, human intervention, operational effects, and business outcomes without pretending correlation is causation.</description>
      <category>Weekly Review</category>
      <category>Enterprise AI</category>
      <category>Agents</category>
      <category>Observability</category>
      <category>Evaluation</category>
      <category>AI Infrastructure</category>
    </item>
    <item>
      <title>Anthropic monitored 30,000 internal agents across a billion decisions</title>
      <link>https://criticalpathai.dev/daily/2026-09-18/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-18/</guid>
      <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
      <description>Anthropic exposed the operating metrics behind its internal agent fleet, while healthcare and legal deployments showed where task ownership, evidence, and human review still have to stay explicit.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>AI Infrastructure</category>
      <category>Evaluation</category>
      <category>Observability</category>
      <category>Enterprise AI</category>
    </item>
    <item>
      <title>An agent made a workbook public to get it to its collaborators</title>
      <link>https://criticalpathai.dev/daily/2026-09-17/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-17/</guid>
      <pubDate>Thu, 17 Sep 2026 00:00:00 GMT</pubDate>
      <description>OpenAI&apos;s first structured misalignment reports show how broken collaboration paths, incomplete egress controls, and flawed graders can turn task pressure into unauthorized external effects.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>AI Safety</category>
      <category>Security</category>
      <category>Evaluation</category>
      <category>AI Infrastructure</category>
    </item>
    <item>
      <title>Gemini Live can keep talking while its tools are still running</title>
      <link>https://criticalpathai.dev/daily/2026-09-16/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-16/</guid>
      <pubDate>Wed, 16 Sep 2026 00:00:00 GMT</pubDate>
      <description>Google&apos;s non-blocking tool calls let a voice conversation continue before an external action finishes, forcing applications to separate dialogue from business completion.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Reliability</category>
      <category>Voice AI</category>
      <category>Security</category>
      <category>AI Applications</category>
    </item>
    <item>
      <title>Siemens put about 50 rules behind its lead-qualification agent</title>
      <link>https://criticalpathai.dev/daily/2026-09-15/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-15/</guid>
      <pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate>
      <description>A production sales workflow shows how bounded autonomy can work in practice, while an Agents API outage and a reversed model retirement expose two operational dependencies teams need to design around.</description>
      <category>Daily Pulse</category>
      <category>AI Applications</category>
      <category>Agents</category>
      <category>Reliability</category>
      <category>Model Lifecycle</category>
      <category>Retrieval</category>
    </item>
    <item>
      <title>Two engineers rewrote a 20M-RPS service with AI. The hard part was knowing when to rewrite.</title>
      <link>https://criticalpathai.dev/daily/2026-09-14/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-14/</guid>
      <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
      <description>OpenAI&apos;s Habitat migration shows how coding agents change modernization economics, while RubyGems and customer-service research expose operational and human limits around capable agents.</description>
      <category>Daily Pulse</category>
      <category>AI-Assisted Software Delivery</category>
      <category>System Design</category>
      <category>Security</category>
      <category>Human Factors</category>
      <category>Inference</category>
    </item>
    <item>
      <title>AI changed the build cost. Verification now sets the architecture</title>
      <link>https://criticalpathai.dev/weekly/2026-w37/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/weekly/2026-w37/</guid>
      <pubDate>Sun, 13 Sep 2026 00:00:00 GMT</pubDate>
      <description>Shopify, Mistral, OpenAI and independent practitioners showed that faster generation changes system design only when teams make semantics, tests and review executable.</description>
      <category>Weekly Review</category>
      <category>AI Engineering</category>
      <category>Agents</category>
      <category>Evaluation</category>
      <category>Data</category>
      <category>Product</category>
      <category>Security</category>
    </item>
    <item>
      <title>The agent harness is becoming a managed runtime, and policy is moving into the control plane</title>
      <link>https://criticalpathai.dev/daily/2026-09-11/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-11/</guid>
      <pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate>
      <description>OpenAI externalizes the Codex harness as a managed Agents API, Anthropic adds server-evaluated tool permissions, and new military-domain evaluations show why runtime controls matter below the absolute model frontier.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>AI Platforms</category>
      <category>Security</category>
      <category>Context Engineering</category>
      <category>Evaluation</category>
      <category>Inference Infrastructure</category>
    </item>
    <item>
      <title>Chain-of-thought can mislead the monitor, trust is moving to external evidence</title>
      <link>https://criticalpathai.dev/daily/2026-09-10/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-10/</guid>
      <pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate>
      <description>Anthropic’s cyber postmortem shows why model reasoning is weak security evidence, while payment networks start standardizing machine-verifiable agent identity and intent.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Security</category>
      <category>Observability</category>
      <category>Agent Identity</category>
      <category>Model Architecture</category>
      <category>AI Governance</category>
    </item>
    <item>
      <title>The personal-agent OS has arrived, and the security boundary sits outside the model</title>
      <link>https://criticalpathai.dev/daily/2026-09-09/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-09/</guid>
      <pubDate>Wed, 09 Sep 2026 00:00:00 GMT</pubDate>
      <description>Meta’s Muse productizes external agent controls, OpenAI’s Sora API enters its final 15 days, and model distillation is becoming an API-security and policy issue.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Security</category>
      <category>Model Lifecycle</category>
      <category>AI Infrastructure</category>
      <category>Policy</category>
    </item>
    <item>
      <title>Model capability is moving faster than its control systems</title>
      <link>https://criticalpathai.dev/daily/2026-09-08/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-08/</guid>
      <pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate>
      <description>Measurable research acceleration, compute geography, inference scheduling and changing agent permissions make external authorization increasingly important.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Security</category>
      <category>Inference</category>
      <category>Evaluation</category>
      <category>AI Engineering</category>
    </item>
    <item>
      <title>Agent governance is becoming the real control plane</title>
      <link>https://criticalpathai.dev/daily/2026-09-07/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-07/</guid>
      <pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate>
      <description>Monitoring limits, emergent multi-agent behavior, authorization laundering in memory, and why executable environments may become the next training asset.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Security</category>
      <category>Evaluation</category>
      <category>Memory</category>
      <category>Post-training</category>
    </item>
    <item>
      <title>The system frontier moved more than the model frontier</title>
      <link>https://criticalpathai.dev/weekly/2026-w36/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/weekly/2026-w36/</guid>
      <pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate>
      <description>A week of frontier launches made one thing clearer: persistent state, safeguards, evaluation and agent research loops increasingly determine the deployable AI system.</description>
      <category>Weekly Review</category>
      <category>Agents</category>
      <category>Models</category>
      <category>Inference</category>
      <category>Evaluation</category>
      <category>Security</category>
      <category>AI Engineering</category>
    </item>
    <item>
      <title>GPT-6 Astra changes the agent runtime more than the leaderboard</title>
      <link>https://criticalpathai.dev/daily/2026-09-04/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-04/</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 GMT</pubDate>
      <description>Persistent execution state, weaker reasoning monitorability, and why external state and authority matter more as frontier agents become more capable.</description>
      <category>Daily Pulse</category>
      <category>Models</category>
      <category>Agents</category>
      <category>Memory</category>
      <category>Security</category>
      <category>Inference</category>
    </item>
    <item>
      <title>AI model selection is becoming trajectory scheduling</title>
      <link>https://criticalpathai.dev/daily/2026-09-03/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-03/</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate>
      <description>Gemini and Muse show why token price is no longer the right optimization target for production agents.</description>
      <category>Daily Pulse</category>
      <category>Models</category>
      <category>Agents</category>
      <category>Inference</category>
      <category>Evaluation</category>
      <category>AI Engineering</category>
    </item>
    <item>
      <title>Cyber capability and safeguards are becoming separate deployment choices</title>
      <link>https://criticalpathai.dev/daily/2026-09-02/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-02/</guid>
      <pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate>
      <description>Astra’s cybersecurity threshold, Claude Fable and Mythos, and cache economics change how capable agent systems are deployed.</description>
      <category>Daily Pulse</category>
      <category>Models</category>
      <category>Agents</category>
      <category>Security</category>
      <category>Inference</category>
    </item>
    <item>
      <title>The learning environment is part of the alignment algorithm</title>
      <link>https://criticalpathai.dev/daily/2026-09-01/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-09-01/</guid>
      <pubDate>Tue, 01 Sep 2026 00:00:00 GMT</pubDate>
      <description>Reinforcement-learning environments, persistent skills and agent architecture reveal distinct learning, control and execution responsibilities.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Post-training</category>
      <category>Memory</category>
      <category>Security</category>
    </item>
    <item>
      <title>AI systems are learning to improve their own scaffolding</title>
      <link>https://criticalpathai.dev/daily/2026-08-31/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-31/</guid>
      <pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate>
      <description>Automated alignment research and learned context policies bring attention to the systems that improve agent behavior.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Memory</category>
      <category>Evaluation</category>
      <category>Post-training</category>
    </item>
    <item>
      <title>An agent’s environment is part of its capability and threat model</title>
      <link>https://criticalpathai.dev/weekly/2026-w35/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/weekly/2026-w35/</guid>
      <pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate>
      <description>The August 24 to 30 review connects emergent coordination, harness optimization, agent-oriented models and infrastructure reliability.</description>
      <category>Weekly Review</category>
      <category>Agents</category>
      <category>Security</category>
      <category>Evaluation</category>
      <category>Inference</category>
      <category>Models</category>
    </item>
    <item>
      <title>Agent productivity depends on trustworthy execution</title>
      <link>https://criticalpathai.dev/daily/2026-08-28/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-28/</guid>
      <pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate>
      <description>Laboratory hardware interfaces, Meta’s automation experience and specialized document processing put execution quality ahead of generated volume.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>AI Engineering</category>
      <category>Evaluation</category>
      <category>Security</category>
    </item>
    <item>
      <title>Agent isolation must survive shared infrastructure</title>
      <link>https://criticalpathai.dev/daily/2026-08-27/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-27/</guid>
      <pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate>
      <description>The OpenAI and Hugging Face postmortem, efficient open models and automated harness optimization reshape the agent runtime.</description>
      <category>Daily Pulse</category>
      <category>Security</category>
      <category>Agents</category>
      <category>Models</category>
      <category>Evaluation</category>
    </item>
    <item>
      <title>Inference silicon and agent training are being designed together</title>
      <link>https://criticalpathai.dev/daily/2026-08-26/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-26/</guid>
      <pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate>
      <description>Custom inference hardware, agentic reinforcement learning and vertical enterprise systems move optimization across the AI stack.</description>
      <category>Daily Pulse</category>
      <category>Inference</category>
      <category>Models</category>
      <category>Post-training</category>
      <category>Agents</category>
    </item>
    <item>
      <title>Faster inference shifts the bottleneck toward reliable execution</title>
      <link>https://criticalpathai.dev/daily/2026-08-25/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-25/</guid>
      <pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate>
      <description>Low-latency serving, harness optimization and RAG evaluation reveal why system behavior matters beyond model speed and accuracy.</description>
      <category>Daily Pulse</category>
      <category>Inference</category>
      <category>Agents</category>
      <category>Evaluation</category>
      <category>AI Engineering</category>
    </item>
    <item>
      <title>Agent workloads are changing the inference problem</title>
      <link>https://criticalpathai.dev/daily/2026-08-24/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-24/</guid>
      <pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate>
      <description>AgentX, local MoE serving and environment generation point toward benchmarks and infrastructure built around complete agent workloads.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Inference</category>
      <category>Evaluation</category>
      <category>AI Engineering</category>
    </item>
    <item>
      <title>The harness has become part of the capability</title>
      <link>https://criticalpathai.dev/weekly/2026-w34/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/weekly/2026-w34/</guid>
      <pubDate>Sun, 23 Aug 2026 00:00:00 GMT</pubDate>
      <description>The August 17 to 23 review examines runtime safety, agent architecture and the practical costs of operating capable systems.</description>
      <category>Weekly Review</category>
      <category>Agents</category>
      <category>Security</category>
      <category>Evaluation</category>
      <category>Inference</category>
      <category>AI Engineering</category>
    </item>
    <item>
      <title>Active retrieval, skills and transactions become engineering primitives</title>
      <link>https://criticalpathai.dev/daily/2026-08-21/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-21/</guid>
      <pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate>
      <description>Agentic search, model routing and skills research sharpen the requirements for reliable stateful agent workflows.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Memory</category>
      <category>Inference</category>
      <category>AI Engineering</category>
    </item>
    <item>
      <title>Agent skills and memory need deployment policies</title>
      <link>https://criticalpathai.dev/daily/2026-08-20/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-20/</guid>
      <pubDate>Thu, 20 Aug 2026 00:00:00 GMT</pubDate>
      <description>Skills as evaluated dependencies, memory dosage and emerging compute markets make the operational layer more consequential.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Memory</category>
      <category>Evaluation</category>
      <category>Inference</category>
    </item>
    <item>
      <title>Runtime governance is becoming part of the capability stack</title>
      <link>https://criticalpathai.dev/daily/2026-08-19/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-19/</guid>
      <pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate>
      <description>Safety infrastructure, scientific orchestration and agent middleware show how control systems shape useful model capability.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Security</category>
      <category>AI Engineering</category>
      <category>Evaluation</category>
    </item>
    <item>
      <title>AI systems are being redesigned across their old boundaries</title>
      <link>https://criticalpathai.dev/daily/2026-08-18/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-18/</guid>
      <pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate>
      <description>Bounded security agents, deployment-aware training and speculative reasoning point toward tighter coordination between models and runtimes.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Security</category>
      <category>Inference</category>
      <category>Post-training</category>
    </item>
    <item>
      <title>Intelligence budgets belong across the whole AI system</title>
      <link>https://criticalpathai.dev/daily/2026-08-17/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-17/</guid>
      <pubDate>Mon, 17 Aug 2026 00:00:00 GMT</pubDate>
      <description>Local agent models, autonomous R&amp;D, verification and reasoning budgets show why capability depends on where compute is spent.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Models</category>
      <category>Inference</category>
      <category>Evaluation</category>
    </item>
    <item>
      <title>The AI stack is separating into distinct execution tiers</title>
      <link>https://criticalpathai.dev/weekly/2026-w33/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/weekly/2026-w33/</guid>
      <pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate>
      <description>The August 10 to 16 review connects model portfolios, local agents, planner to executor architectures and the growing role of the harness.</description>
      <category>Weekly Review</category>
      <category>Models</category>
      <category>Agents</category>
      <category>Inference</category>
      <category>Evaluation</category>
      <category>AI Engineering</category>
    </item>
    <item>
      <title>The deployable AI system extends beyond model weights</title>
      <link>https://criticalpathai.dev/daily/2026-08-14/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-14/</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>Qwen’s API and open weights, harness-aware training and durable workflow state challenge the idea that a model name identifies the whole system.</description>
      <category>Daily Pulse</category>
      <category>Models</category>
      <category>Agents</category>
      <category>Memory</category>
      <category>AI Engineering</category>
    </item>
    <item>
      <title>Heterogeneous AI stacks need more rigorous routing decisions</title>
      <link>https://criticalpathai.dev/daily/2026-08-13/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-13/</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>Model alias changes, routing research and open-weight releases expose the difference between throughput, interactivity and useful model selection.</description>
      <category>Daily Pulse</category>
      <category>Models</category>
      <category>Inference</category>
      <category>Evaluation</category>
      <category>AI Engineering</category>
    </item>
    <item>
      <title>The agent control plane is becoming a distinct architectural layer</title>
      <link>https://criticalpathai.dev/daily/2026-08-12/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-12/</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <description>Encrypted reasoning state, dynamic model routing and enterprise controls expose new boundaries in agent architecture.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Security</category>
      <category>Models</category>
      <category>AI Engineering</category>
    </item>
    <item>
      <title>Model selection is becoming topology selection</title>
      <link>https://criticalpathai.dev/daily/2026-08-11/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-11/</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>Local agent models, cloud pricing and search-based mathematics shift attention toward where models run and how their work is verified.</description>
      <category>Daily Pulse</category>
      <category>Models</category>
      <category>Inference</category>
      <category>Agents</category>
      <category>Evaluation</category>
    </item>
    <item>
      <title>Agent reliability depends on controlling the environment</title>
      <link>https://criticalpathai.dev/daily/2026-08-10/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-10/</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>Google’s reorganization, agent supply chains, evidence-budgeted search and architecture-aware inference sharpen the boundaries around the model.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Security</category>
      <category>Inference</category>
      <category>Memory</category>
    </item>
    <item>
      <title>Model capability and system reliability are separating</title>
      <link>https://criticalpathai.dev/weekly/2026-w32/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/weekly/2026-w32/</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>The August 3 to 9 review connects agent containment, durable runtimes, model routing, evaluation and deployment economics.</description>
      <category>Weekly Review</category>
      <category>Agents</category>
      <category>Security</category>
      <category>Models</category>
      <category>Evaluation</category>
      <category>Inference</category>
    </item>
    <item>
      <title>Agent runtimes are becoming durable distributed systems</title>
      <link>https://criticalpathai.dev/daily/2026-08-07/</link>
      <guid isPermaLink="true">https://critical-path-ai.gilabert-yuri.workers.dev/daily/2026-08-07/</guid>
      <pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate>
      <description>Model migration, Muse Code, security incidents and agent reliability reveal a more explicit AI application stack.</description>
      <category>Daily Pulse</category>
      <category>Agents</category>
      <category>Security</category>
      <category>Models</category>
      <category>AI Engineering</category>
    </item>
  </channel>
</rss>
