In this article
Today’s strongest signal is that agent workloads are starting to reshape the infrastructure beneath them. Traditional inference benchmarks increasingly look artificial once real coding agents carry 100K+ token histories, reuse enormous prefixes, spawn subagents and repeatedly mutate their context. Meanwhile, open-source serving research is treating CPU RAM, PCIe, GPU memory, KV state and MoE experts as one dynamic memory hierarchy; environment-generation research is turning executable worlds themselves into a trainable artifact; and enterprise usage data suggests the economic default is moving away from “always use the smartest model.”
AI infrastructure is becoming trajectory-aware rather than request-aware.
That distinction should influence model gateways, inference benchmarks, cache architecture and agent evaluation.
AgentX may be a more important inference benchmark than another tokens/sec leaderboard
SemiAnalysis released AgentX 1.0 / InferenceXv3 today, an Apache-2.0 benchmark specifically for multi-turn agentic coding inference at context lengths up to 1M tokens. The headline is not another hardware ranking; the important part is that the workload shape comes from real agent traces rather than synthetic 8K input → 1K output requests. (SemiAnalysis)
The v1.0 dataset contains 393 Claude Code sessions with at least 20 requests each, selected from an internal opt-in corpus. It is a coding-workload sample, not a representative sample of every agent application. Median input per request is roughly 142K tokens, median output only around 444 tokens, and 44% of sessions involve subagents. The benchmark explicitly models repeated turns, growing histories, concurrent child agents and warm cache state instead of treating every request as independent. (InferenceX) For these coding workloads, prefill and decode speed need to be evaluated alongside prefix reuse, cache residency, context mutation, scheduler fairness, subagent concurrency and tail latency.
When 95%+ of a giant prefix can potentially be reused, recomputing context because your cache policy evicted the wrong blocks can matter more than a modestly faster attention kernel.
The nice part is that AgentX is already exposing real engineering consequences. Work driven by the benchmark improved vLLM’s handling of long-lived prefixes to above 95% KV-cache hit rates in tested 1M-context workloads. For DeepSeek-V4 under HBM pressure, a hybrid CPU-KV-offload strategy reportedly delivered 81.7% higher output throughput and 46.6% lower mean end-to-end latency than recomputing evicted prefixes. (InferenceX) That tells us something important about agent-serving economics. For a stateless chatbot, losing a cache entry is annoying. For a 300K-token coding trajectory with twelve turns and several subagents, losing the wrong state can trigger enormous redundant prefill. I would therefore start treating trajectory cost as the unit of inference benchmarking: total accelerator time + CPU/offload cost + cache churn + wall time
per completed agent trajectory, rather than benchmarking isolated API calls. And the benchmark should preserve workload topology. Ten concurrent independent users and one parent agent spawning nine workers may have identical request concurrency but very different prefix-sharing opportunities. This also creates a healthy feedback loop between application traces and serving engines: real agents → workload benchmark → runtime optimization → cheaper/faster agents → new workload shapes. That is how mature infrastructure ecosystems evolve. For inference capacity planning, study the workload assumptions before comparing accelerator rankings. (SemiAnalysis)
FreeToken shows what local MoE serving looks like when you stop pretending RAM is just “slow VRAM”
A strong open-source project from the research pass is FreeToken, developed by a group including researchers associated with Berkeley, Stanford systems work and the broader SGLang ecosystem. It targets something increasingly relevant: running very large sparse open-weight models on ordinary personal hardware. The code is Apache-2.0. (arXiv) The interesting design choice is that FreeToken does not use a fixed “GPU model with CPU offloading” abstraction. Instead, it treats: GPU compute + GPU memory + CPU compute + host RAM + PCIe as one heterogeneous execution system.
For MoE models this makes sense because inference has a peculiar resource profile. Only a fraction of experts execute for each token, yet the complete expert set may be hundreds of gigabytes. Keeping everything in VRAM is impossible on consumer hardware; blindly fetching every cache miss over PCIe is painfully slow.
FreeToken dynamically chooses whether a missing expert should be transferred to the GPU or executed directly on the CPU. Its q* policy benchmarks the actual host-memory and PCIe bandwidth of the machine and splits misses accordingly. Hot experts remain in a GPU-side LRU cache. (GitHub) There is another subtle complication. MoE is sparse during decode, but prefill can look much less sparse operationally: across tens of thousands of prompt tokens, routing may touch most experts. Streaming the whole layer through a tiny expert cache therefore destroys locality. FreeToken instead uses layer-level streaming/double buffering for that phase and switches execution strategies as the workload changes. (arXiv)
Its memory manager is also agent-aware. As an agent conversation grows, KV state consumes increasing VRAM. FreeToken can shrink the expert cache and expand KV capacity dynamically instead of fixing the split at server startup. It also implements semantic anchor checkpoints so context edits caused by tool calls or reasoning-block changes can reuse earlier recurrent/KV state rather than always recomputing from token zero. (GitHub)
The authors report running a 35B model on an 8GB laptop GPU, a 284B model on a gaming desktop and GLM-5.2 753B on a single workstation GPU, using the host CPU and RAM as part of the serving system. Those claims do not mean the full weights fit in GPU memory. Those are author benchmarks and need independent reproduction, so I would not treat the throughput claims as established yet. (arXiv) The architecture treats local inference as a memory-placement problem across CPU RAM, GPU VRAM and their interconnect, managed by the scheduler. Quantization is only one part of that system.
That matters for sovereign/local AI because sparse models can have enormous total capacity while exposing a much smaller active computational footprint. A sufficiently intelligent runtime can exploit that asymmetry.
I would not replace a production serving stack with it yet, but the paper is worth studying because several of its ideas (elastic expert/KV allocation, bandwidth-aware execution and semantic-state reuse)will likely recur elsewhere. (arXiv)
AgentMercury: generating the world may become more important than generating more tasks
Today’s Hugging Face research feed surfaced AgentMercury, published August 21, and I think it captures an important next step in agent training. Most synthetic-agent benchmarks start with a task:
process this invoice.
The benchmark designer then builds just enough environment around that task for the agent to execute it. AgentMercury reverses the process. Given a high-level scenario, it synthesizes a persistent executable world containing entities, services, tools, mutable state and executable cross-service invariants. Tasks then emerge from that world. The authors generated 4,783 environments spanning 14 industries and 50 countries. (arXiv) That distinction is important. A task-specific environment teaches: how to solve this task. A coherent world can teach: how actions alter a system in which many possible tasks exist.
The authors used those environments for RL and report that Qwen3.5-4B moved from 12.3 to 15.7 on EnterpriseOps-GYM and from 45.9 to 56.0 on AIME26 despite the environments not being generated specifically for those benchmarks. More intriguingly, environment construction itself became trainable: fine-tuning Qwen3.5-35B-A3B on construction traces reportedly increased successful executable-world generation from 3.3% to 83.3% on held-out scenarios. These are authors’ results and need reproduction. (arXiv) The architectural idea is stronger than the benchmark numbers. Imagine creating a synthetic enterprise containing:
- CRM customers;
- orders;
- contracts;
- support cases;
- inventory;
- billing;
- identity;
and deterministic invariants such as: refund total ≤ captured payment or terminated employee cannot access production. Agents can then practice tens of thousands of workflows inside that environment while the invariants provide machine-verifiable reward. This connects directly to a useful thesis from Latent Space this weekend: synthetic generation tends to become useful when verification catches up. Synthetic text became useful once filtering improved; RL environments became useful where rewards/tests were trustworthy; scientific simulations become useful where physical or formal validation exists. (Latent.Space) The danger is obviously simulator overfitting. If your synthetic enterprise misunderstands tax rules, authorization boundaries or customer behavior, the RL agent may become extremely competent at exploiting a world that does not exist. So the difficult part is not generating environments.
It is defining trustworthy world invariants and fidelity tests. For production agent teams, that suggests a future artifact hierarchy: production system → executable behavioral specification → training/evaluation world rather than building every benchmark manually. I would track this family of work closely. (arXiv)
Adjacent research signal: specialization is still a powerful substitute for scale
Also worth watching is Faraday, a 27B “AI scientist” model trained specifically on research-paper replication. Its creators built a scalable replication task environment and post-trained the model to orchestrate coding agents as tools. They report Faraday outperforming Claude Opus 4.8 and GPT-5.5 on their held-out replication tasks. (arXiv)
The large caveat is obvious: the team built both the task space and the model, so independent evaluation is essential.
The interesting signal is narrower: a 27B specialized agent can potentially outperform much stronger general models inside a well-designed domain environment. That’s another reason not to equate agent capability with base-model scale.
Enterprise users may finally be refusing the “always buy the frontier” default
The most strategically relevant business signal today comes from Ramp’s August AI Index, amplified by Financial Times reporting over the weekend.
Ramp’s July token-management data shows Anthropic Fable 5 accounting for only 6% of Anthropic tokens and 11.4% of model spend, despite being Anthropic’s highest-performing and most expensive offering. OpenAI’s GPT-5.6 Sol represented 25% of OpenAI tokens and 23% of spend. Ramp explicitly interprets Fable as evidence of an upper bound at which incremental model capability is no longer worth the price for many enterprise workloads. Its token-management sample skews more technical than Ramp’s broader customer base, so treat the exact percentages accordingly. (Ramp)
The FT reports that cheaper Anthropic models have continued gaining share relative to Fable, reinforcing that the pattern persisted beyond the initial launch window. (Financial Times) This is more architecturally interesting than it sounds. Until recently, a surprisingly defensible strategy was:
Spend minimal effort optimizing model assignment because next month’s frontier model will probably dominate everything anyway.
Drew Breunig captured the change neatly in a comment highlighted by Simon Willison: once the frontier model becomes sufficiently expensive, the work of deciding what goes where suddenly matters. (Simon Willison’s Weblog) This favors using a workhorse model by default and escalating when uncertainty, consequence or task difficulty justifies the frontier model’s cost. Crucially, this does not imply enterprises are abandoning frontier providers. Ramp’s earlier data shows companies using model-serving/open-model platforms are often heavier users of OpenAI and Anthropic as well; the association is consistent with complementary usage, but does not establish that cheaper models caused total workload volume to increase. (Ramp) So the likely future is not one winning model. It is portfolio economics.
The next model-gateway advantage may come less from finding the objectively “best” model and more from deciding where the marginal intelligence of the expensive model changes the final outcome.
From the technical feeds
SemiAnalysis / InferenceX. AgentX. This is today’s best technical-feed contribution. The value is not the CUDA-versus-AMD horse race; it is a genuinely more realistic inference workload: long-lived sessions, >100K-token average inputs, subagents and extremely high prefix-cache reuse. Read the methodology if you design inference or benchmark vendors. (SemiAnalysis)
Latent Space. “Why Simulation is taking over.” The useful thesis is that synthetic components become load-bearing not when generators suddenly stop hallucinating, but when verification mechanisms make their mistakes cheap to detect. AgentMercury is a nice concrete example of the same principle: executable invariants matter more than the ability to generate another plausible-looking enterprise world. (Latent.Space)
Simon Willison / Drew Breunig, the end of the frontier-model free lunch. Very short, but worth retaining: when frontier capability is automatically affordable, harness optimization and routing can look premature; once the capability/cost curve spreads out, context strategy and workload placement become core engineering work. Ramp’s adoption data now provides empirical support for that intuition. (Simon Willison’s Weblog)
Ramp’s spending data represents its U.S. business customers; the token-level subset is more technical still. It should not be generalized to the entire model market. (Ramp)
Stay Sharp: Two caches, two locality problems
AgentX and FreeToken expose a useful distinction that becomes increasingly important for MoE agents. A long-running agent may simultaneously depend on KV/prefix locality and expert locality. The KV cache stores intermediate attention state for the conversation prefix. Its reuse pattern is primarily temporal across turns: turn 1 prefix → reused by turn 2 → reused by turn 3. If that state remains resident, the model can avoid re-prefilling 200K tokens. Lose it, and you may pay a huge recomputation penalty. An MoE expert cache solves a different problem. Hundreds of experts exist, but only some are selected on each token. Its locality is primarily routing locality across nearby tokens and requests: expert 37 becomes hot → keep its weights in VRAM → avoid repeatedly pulling them across PCIe.
Those caches compete for the same GPU memory, but have different trust properties. Model expert weights may be shared across tenants; cached prompt state can encode private data and needs authorization, isolation and deletion controls. A high cache-hit rate does not justify crossing those boundaries. Early in an agent session, a small KV allocation and larger expert cache may be optimal. After a 400K-token trajectory develops, retaining more KV state and fewer expert weights may be worthwhile because reconstructing the prefix has become more expensive. This is precisely why FreeToken’s dynamic VRAM split is interesting. (GitHub) And it explains why the serving scheduler is gradually becoming a memory-hierarchy manager rather than merely a request queue. Cache-hit rate alone cannot decide the allocation. The scheduler needs to retain the state whose expected reconstruction cost makes residency most valuable.
That is classical systems thinking resurfacing inside LLM inference.
Worth Your Time
- AgentX / InferenceXv3: read the benchmark methodology before the hardware conclusions; its workload model is likely to influence how agentic inference gets measured. AgentX announcement
- FreeToken: read the paper plus the open-source runtime, especially the
q*hybrid execution and elastic KV/expert-memory management sections. FreeToken paper FreeToken repository - AgentMercury: strongest research read today if your interests include agent training, evaluation or synthetic enterprise environments. AgentMercury paper
Today’s architectural takeaway: the request is becoming the wrong unit of abstraction. Agent systems have trajectories, shared prefixes, subagents, evolving working sets, specialist policies, persistent worlds and differentiated model budgets. Inference engines, routers and evaluation systems that understand those structures can produce larger gains than another isolated optimization to the model call itself.