The publication index
Topics
From agents to inference, follow the recurring themes across Daily Pulses and Weekly Reviews.
Agents
A typed evaluator matched five labels across 500 repeated decisions
Evaluation
OpenAI's four-model shutdown hides three migrations behind one replacement
Security
An agent made a workbook public to get it to its collaborators
Inference
Two engineers rewrote a 20M-RPS service with AI. The hard part was knowing when to rewrite.
AI Engineering
Gemini 4 Argon's coding score needs a benchmark audit
Models
The system frontier moved more than the model frontier
Memory
Agent governance is becoming the real control plane
AI Infrastructure
A typed evaluator matched five labels across 500 repeated decisions
Post-training
Agent governance is becoming the real control plane
AI Governance
GPT-Rosalind starts billing today. Price the whole scientific run
Model Lifecycle
Gemini 4 Argon's coding score needs a benchmark audit
Agent Security
Claude moved Cowork's execution boundary. Rebuild the trust map
AI Architecture
NVIDIA put the agent watchdog outside the host
Observability
AI teams are starting to measure work, not usage
Reliability
Changing an agent’s tools now has a cache bill and a replay contract
Agent Architecture
OpenAI gave its always-on agent a read-only research mode
Agent Evaluation
Ironclad turned 11 workflows into agent tests. The rubric is the release artifact
Agent Identity
Anthropic now bills some refusals with zero output tokens
AI Applications
Gemini Live can keep talking while its tools are still running
AI Economics
OpenAI's four-model shutdown hides three migrations behind one replacement
Context Engineering
Kimi-K3 solved 94% of workflows once. Only 13% survived every run
Enterprise AI
AI teams are starting to measure work, not usage
MCP
GPT-Rosalind starts billing today. Price the whole scientific run
Model Migration
NVIDIA put the agent watchdog outside the host
Reliability Engineering
Kimi-K3 solved 94% of workflows once. Only 13% survived every run
Agent Harnesses
Kimi-K3 solved 94% of workflows once. Only 13% survived every run
Agent Operations
Claude moved Cowork's execution boundary. Rebuild the trust map
Agentic Commerce
Meta’s agent handed calls to people. Delegation needs its own control plane
AI Agents
Changing an agent’s tools now has a cache bill and a replay contract
AI for Science
Kimi-K3 solved 94% of workflows once. Only 13% survived every run
AI Memory
Google says user devices will hold the keys to persistent AI memory
AI Platforms
The agent harness is becoming a managed runtime, and policy is moving into the control plane
AI Safety
An agent made a workbook public to get it to its collaborators
AI-Assisted Software Delivery
Two engineers rewrote a 20M-RPS service with AI. The hard part was knowing when to rewrite.
API Economics
Anthropic now bills some refusals with zero output tokens
API Migration
OpenAI's four-model shutdown hides three migrations behind one replacement
Autonomous Agents
OpenAI gave its always-on agent a read-only research mode
Benchmarking
Gemini 4 Argon's coding score needs a benchmark audit
Business Workflow Automation
Ironclad turned 11 workflows into agent tests. The rubric is the release artifact
Causal Attribution
Fin's 8.6% hard-resolution lift needs a feedback denominator
Cloud Sandboxes
Claude moved Cowork's execution boundary. Rebuild the trust map
Computer Use
Ironclad turned 11 workflows into agent tests. The rubric is the release artifact
Contract Operations
Ironclad turned 11 workflows into agent tests. The rubric is the release artifact
Customer Support AI
Fin's 8.6% hard-resolution lift needs a feedback denominator
Data
AI changed the build cost. Verification now sets the architecture
Data Governance
Claude moved Cowork's execution boundary. Rebuild the trust map
Decision Systems
OpenAI gave its always-on agent a read-only research mode
Desktop Integration
Claude moved Cowork's execution boundary. Rebuild the trust map
Distillation
Gemini 4 Argon's coding score needs a benchmark audit
Failure Taxonomy
Anthropic now bills some refusals with zero output tokens
Feedback Systems
Fin's 8.6% hard-resolution lift needs a feedback denominator
Human Factors
Two engineers rewrote a 20M-RPS service with AI. The hard part was knowing when to rewrite.
Human Oversight
Meta’s agent handed calls to people. Delegation needs its own control plane
Identity and Authorization
Meta’s agent handed calls to people. Delegation needs its own control plane
Inference Infrastructure
The agent harness is becoming a managed runtime, and policy is moving into the control plane
Information Retrieval
Claude moved Cowork's execution boundary. Rebuild the trust map
Key Management
Google says user devices will hold the keys to persistent AI memory
Local Inference
Anthropic now bills some refusals with zero output tokens
MLOps
A typed evaluator matched five labels across 500 repeated decisions
Model Architecture
Chain-of-thought can mislead the monitor, trust is moving to external evidence
Model Deprecations
OpenAI gave its always-on agent a read-only research mode
Model Economics
GPT-Rosalind starts billing today. Price the whole scientific run
Model Evaluation
Gemini 4 Argon's coding score needs a benchmark audit
Model Post-Training
Fin's 8.6% hard-resolution lift needs a feedback denominator
Model Testing
Ironclad turned 11 workflows into agent tests. The rubric is the release artifact
Policy
The personal-agent OS has arrived, and the security boundary sits outside the model
Policy Enforcement
NVIDIA put the agent watchdog outside the host
Privacy Engineering
Google says user devices will hold the keys to persistent AI memory
Product
AI changed the build cost. Verification now sets the architecture
Production Evaluation
Fin's 8.6% hard-resolution lift needs a feedback denominator
Prompt Caching
Changing an agent’s tools now has a cache bill and a replay contract
Research Reproducibility
GPT-Rosalind starts billing today. Price the whole scientific run
Retrieval
Siemens put about 50 rules behind its lead-qualification agent
Runtime Isolation
NVIDIA put the agent watchdog outside the host
Scientific AI
GPT-Rosalind starts billing today. Price the whole scientific run
Software Engineering
Gemini 4 Argon's coding score needs a benchmark audit
Stateful Testing
Ironclad turned 11 workflows into agent tests. The rubric is the release artifact
System Design
Two engineers rewrote a 20M-RPS service with AI. The hard part was knowing when to rewrite.
Trusted Execution
Google says user devices will hold the keys to persistent AI memory
Voice AI
Gemini Live can keep talking while its tools are still running