Kimi-K3 solved 94% of workflows once. Only 13% survived every run
New evidence separates occasional agent capability from dependable operation, while applied teams show where task choice, model training, and runtime retrieval each belong.
The bigger picture · 9 editions
A deeper synthesis of models, agents, infrastructure and research. We focus on the patterns that should change architectural assumptions.
New evidence separates occasional agent capability from dependable operation, while applied teams show where task choice, model training, and runtime retrieval each belong.
Muse exposed a specific agent boundary: authenticating the visible service does not reveal who performed an action, what authority crossed the handoff, or who owes the user a remedy.
OpenAI and Anthropic exposed richer task-level metrics this week. The architectural opportunity is a task ledger that connects model activity, human intervention, operational effects, and business outcomes without pretending correlation is causation.
Shopify, Mistral, OpenAI and independent practitioners showed that faster generation changes system design only when teams make semantics, tests and review executable.
A week of frontier launches made one thing clearer: persistent state, safeguards, evaluation and agent research loops increasingly determine the deployable AI system.
The August 24 to 30 review connects emergent coordination, harness optimization, agent-oriented models and infrastructure reliability.
The August 17 to 23 review examines runtime safety, agent architecture and the practical costs of operating capable systems.
The August 10 to 16 review connects model portfolios, local agents, planner to executor architectures and the growing role of the harness.
The August 3 to 9 review connects agent containment, durable runtimes, model routing, evaluation and deployment economics.