Anthropic’s test agent filed a false homicide tip. Evaluation needs a blast radius
An agent test created a real external effect. Google, Uber, and new judge benchmarks show how to separate observation, diagnosis, policy, and permission to act.
Topic · 1 article
The latest analysis on LLM-as-a-Judge, across Daily Pulses and Weekly Reviews.
An agent test created a real external effect. Google, Uber, and new judge benchmarks show how to separate observation, diagnosis, policy, and permission to act.