In this article 8 sections

Today’s developments connect inference architecture, post-training in agent environments, vertical enterprise systems and test-time reasoning.

The new systems expose trade-offs that component-level benchmarks alone cannot settle.

OpenAI’s Jalapeño results show hardware, memory, network, scheduling and model workloads being co-designed as one inference system. IBM Granite 4.2 trains directly inside real agent environments instead of hoping generic reasoning transfers cleanly. Google’s new legal and financial AI products treat domain skills, permission-bound connectors and governance as part of the agent, not surrounding integrations. And fresh decoding research suggests that when reasoning goes wrong, regenerating the entire trajectory may be wasteful; repair the damaged section instead.


OpenAI Jalapeño: custom inference silicon becomes strategically credible

OpenAI published the first measured results for Jalapeño, its first custom inference ASIC, on August 25. Across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, OpenAI reports 1.5 to 1.9× more useful AI work per watt at peak throughput, 1.7 to 3.6× lower end-to-end latency, and 2.1 to 4.1× higher performance on highly interactive operating points versus its tested comparison systems. The per-watt comparisons use published accelerator power ratings, not measured whole-rack wall power. Jalapeño is rated at 700 W, although OpenAI says sustained chip power stayed at or below 550 W in those tests. (OpenAI) The numbers matter, but the architectural decision behind them matters more.

Jalapeño is explicitly designed around the observation that LLM inference is not one workload. Prefill is relatively compute-heavy; decode is more bandwidth-sensitive; KV state introduces locality constraints; inter-chip communication can dominate once tensors and state start moving. OpenAI therefore co-designed chip, memory, fabric, software and rack around keeping state local and allowing the same resources to service changing inference phases. (OpenAI) This is interesting because it goes somewhat against one dominant inference trend: prefill/decode disaggregation. A clean P/D split looks attractive: prefill pool → KV transfer → decode pool because each fleet can be optimized for one phase.

SemiAnalysis’ technical analysis says OpenAI instead chose a more fungible homogeneous system. The rationale is workload variance: input length, output length, cache hit rate, speculative-decoding acceptance, concurrency and latency targets all change continuously. A fixed P/D ratio can therefore strand expensive capacity when demand shifts. It also forces large KV state to cross a network boundary exactly when decode wants low latency. (SemiAnalysis) This gives us a useful infrastructure rule: phase specialization improves local efficiency; fungibility can improve fleet efficiency. Neither architecture is universally better. A hyperscale workload with stable proportions may justify deeply specialized pools. A volatile mixture of interactive agents, batch inference, huge contexts and changing model architectures may favor flexible hardware even if some units are temporarily underutilized within an individual phase.

There are two important caveats to the headline. First, the published OpenAI benchmark is mostly nominal 8K-input / 1K-output, single-turn InferenceX, not the multi-turn long-context AgentX workloads SemiAnalysis introduced earlier this week. SemiAnalysis explicitly says it has not seen AgentX results, where prefix caching, routing, state offload and cache management become much more important. (SemiAnalysis)

Second, comparisons against GB200/GB300 are useful but not the whole competitive picture. SemiAnalysis argues that Vera Rubin is the more appropriate generation-level comparison because Jalapeño also targets HBM4-era deployment. Jalapeño remains promising in the public operating points SemiAnalysis examined, but the production competition will be against newer systems than yesterday’s Blackwell installations. (SemiAnalysis)

OpenAI says it plans to begin deploying Jalapeño internally by the end of 2026, while Gen 2 is already deep in development and Gen 3 is being designed. (OpenAI)

Why this matters architecturally

An inference platform combines the model with its precision, kernels, cache policy, fabric, scheduler, workload distribution and latency SLO. A model-to-accelerator assignment alone does not specify the deployable system. It also reinforces a wider strategic trend: frontier labs are acquiring enough internal workload data to design the hardware around the real shape of their agents, rather than accepting whatever abstraction the general GPU market provides. I would read the architecture and benchmark methodology, not just the “beats NVIDIA” framing.


IBM Granite 4.2: agentic RL is becoming a concrete training stage, not a marketing adjective

IBM released Granite 4.2 on August 25 in 3B, 8B and 30B dense models under Apache 2.0. All three add explicit thinking/non-thinking modes, low-effort reasoning and native tool calling. The 8B and 30B models receive an additional agentic RL phase targeting software engineering, terminal operation and search-driven workflows. (IBM Research) The most interesting part of the release is how explicit IBM is about the training environments. The agentic stages execute the model inside real sandboxed environments, with multi-turn tool use and sparse outcome rewards. The curriculum proceeds through software-engineering tasks, terminal tasks and search. The terminal stage can involve up to 64 environment interactions in a rollout. IBM’s RL infrastructure exposes the same interface whether the verifier is a simple deterministic checker or a full SWE sandbox. (Hugging Face)

Training now includes the agent loop itself: acting, observing how the environment changes, recovering from errors and having an external verifier score the final outcome. Supervised tool-call examples cover only part of that behavior. That strengthens an argument from the past two weeks: model and harness compatibility will increasingly be behavioral rather than syntactic. An OpenAI-compatible endpoint can reduce schema adaptation, but tool-call, error and streaming semantics still need contract tests. It does not mean two models learned under equivalent execution environments.

Granite 4.2 also illustrates how much capability can now be packed into a relatively modest deployment tier. IBM pre-trained the family on roughly 15T tokens, includes long-context training up to 512K, and explicitly positions the smaller models for high-throughput agent work while reserving the 30B tier for harder reasoning and coding tasks. (Hugging Face)

I would not over-index on IBM’s benchmark charts until independent results accumulate. The more durable signal is the training recipe.

The interesting systems detail: asynchronous RL

IBM runs its GRPO pipeline asynchronously. Generation workers continuously produce trajectories into a buffer while separate trainer workers consume batches, update the policy, and stream new weights back without stopping the rollout fleet. (Hugging Face)

That is particularly important for agentic RL because environment interaction is expensive. Waiting for shell commands, repositories, search systems or verifiers means rollout time can dwarf the optimizer step. Keeping the rollout fleet busy therefore becomes an infrastructure problem independent of the RL algorithm itself. Granite 4.2 may or may not become the model you deploy; IBM’s agentic-RL architecture is worth understanding regardless.


Google launched Gemini Enterprise for Legal and Gemini Enterprise for Financial Services on August 25. The legal release is particularly explicit about the architecture: Google says general-purpose model intelligence is necessary but “nowhere near sufficient” for production legal work. Its solution combines domain skills, secure data/system connectors, task-completing agents, partner integrations and a governance layer underneath all of them. (Google Cloud)

The “skills” are reusable packages of instructions and context intended to encode firm-specific playbooks, citation requirements and operating style, for example contract review/redlining, legal research or regulatory monitoring. (Google Cloud)

The connectors are even more important. Google is using secure MCP-based connections into systems such as document repositories and enterprise applications while inheriting the underlying platform’s user-level permissions instead of flattening access for the agent. (Google Cloud) Propagating the user’s identity through a permission-aware connector lets the source system enforce access. Giving an agent a broadly privileged service account and asking the model to decide what the user may see puts authorization in the wrong component. Reuters reports that several major law firms are involved in shaping the legal product, and the broader rollout fits a growing competition among Google, Anthropic, Thomson Reuters and specialist legal-AI vendors. (Reuters)

The deeper implication

A vertical application can combine a strong general model with domain procedures, authoritative data, source-system permissions, specialist tools or models, governance and evaluation. A separately trained foundation model is not always required. That is encouraging for enterprise architecture because those components can evolve independently. You can replace the general model without rebuilding the firm’s playbook. You can update regulatory knowledge without retraining the model. You can enforce matter-level permissions in the source system rather than hoping the prompt correctly remembers the ethical wall. This is also another validation of the “skills as behavioral dependencies” theme from last week: Google is now productizing almost exactly that concept. Use this as a reference for how domain agents combine governed data, reusable procedures and access controls.


Research watch: stop regenerating an entire reasoning trajectory when only the suffix is broken

A fresh August 25 paper, Selective Regenerative Decoding (SRD), attacks an inefficiency in Best-of-N and rejection-sampling approaches to test-time reasoning. Most current inference-time selection schemes treat each candidate trajectory atomically: good → keep and bad → discard. But many failed trajectories are not bad from token one. They contain a useful reasoning prefix and then take a wrong turn.

SRD instead classifies candidates into keep, discard, or refine, and for the refine group regenerates only the degraded suffix while retaining the useful prefix. (ArcXiv)

The authors report a theoretical 1.28 to 1.36× sample-efficiency gain over rejection sampling under their assumptions, plus matched Best-of-N accuracy with substantially fewer generated tokens across MATH500, GPQA Diamond, HotpotQA and AlpacaEval. Those are paper-reported results and will need independent replication. (ArcXiv) The idea is more broadly useful than the exact gains. There is a spectrum: regenerate answer → regenerate trajectory → repair from divergence point. The third becomes increasingly attractive as trajectories grow longer. Imagine an architecture-planning agent that correctly establishes:

  1. workload characteristics;
  2. data sensitivity;
  3. availability requirements;
  4. storage semantics;

but then incorrectly assumes that service X supports a needed consistency mode. Naïve Best-of-N throws away the entire plan. A repair-oriented runtime identifies the critical false predicate and asks the model to resume from the last still-valid checkpoint. This is essentially checkpoint/recovery applied to inference-time reasoning. Before replaying an effectful trajectory, also check which external actions already committed; retaining a reasoning prefix cannot undo them. The hard part is reliable localization: a verifier must distinguish “the conclusion is wrong” from “this particular step is where the reasoning became invalid.” But if that signal becomes good enough, trajectory repair could be far more economical than generating five complete alternatives. I would particularly watch whether the idea transfers from benchmark reasoning into tool-using agent trajectories.


API lifecycle watch: OpenAI Assistants API shuts down today

One transition worth mentioning because the date is concrete: OpenAI’s Assistants API reaches its announced shutdown date on August 26, 2026. OpenAI deprecated it after the Responses API reached feature parity and recommends Responses for new and migrated agent integrations. Responses folds in persistent conversations plus hosted tools such as computer use, MCP and code execution, and supports multi-step model/tool workflows through a more composable API. (OpenAI Developers) This is more than an API rename. Assistants encoded a relatively heavyweight hosted object: Assistant = model + instructions + tools + threads. Responses shifts the center of gravity toward composable execution, where applications can own their model/instruction/tool configuration while conversation state is represented separately.

That aligns better with the architecture we have been seeing everywhere else: agent identity/configuration, conversation state and runtime execution should not necessarily be one monolithic resource.

If any production integration is still calling Assistants endpoints, today is the hard deadline rather than another deprecation reminder. (OpenAI Developers)


From the technical feeds

SemiAnalysis. OpenAI Jalapeño. This is the strongest feed item today because SemiAnalysis was physically in the lab validating the InferenceX runs and adds exactly the nuance the primary announcement lacks: no AgentX results yet, the Blackwell-vs-Jalapeño comparison is generationally imperfect, and OpenAI’s decision not to use fixed prefill/decode pools is an explicit fleet-utilization trade-off rather than an omission. Read it after the OpenAI post. (SemiAnalysis)

Daily Dose of Data Science, the “join problem” in agent retrieval. Yesterday’s GTM-agent walkthrough is partly product-led, but it contains an excellent systems idea. Many agent questions are not retrieval problems; they are relational joins across heterogeneous entities. A search API that returns snippets forces the agent to fetch, parse and reconstruct structured records before it can even perform the join. Returning full typed records makes the LLM’s job a merge over reliable structures rather than information reconstruction from fragments. (Daily Dose of Data Science)

When a task relates customers, invoices, contracts and support history (or services, deployments, incidents and owners)expose those structured entities and perform deterministic joins where the semantics are known. Four independent text-search tools would leave the model to reconstruct relationships the system already knows.

Latent Space. AI Engineering becomes an explicit discipline. Its August 25 issue highlights DeepLearning.AI’s renewed emphasis on AI Engineering after analysis of more than 10,000 job postings and industry interviews. The four themes Latent Space surfaces, building/deploying AI applications with evaluation loops, strong software fundamentals, operating coding agents, and understanding the business/build context, are not novel individually, but the institutional shift is useful evidence that “AI engineering” is settling into a systems discipline rather than prompt specialization. (Latent.Space)

I excluded ByteByteGo’s fresh encrypted-reasoning-trace article from today’s picks because we already covered that attack and its architectural implications in depth; repeating it just because a new explainer appeared would violate the purpose of the feed layer. (ByteByteGo)


Stay Sharp: Why asynchronous RL is useful, and why policy staleness matters

IBM’s Granite pipeline gives us a good reason to refresh a subtle systems problem in RL. In a simple synchronous RL loop: policy θₜ → generate rollouts → score → update θₜ₊₁ → generate again. The rollout fleet waits during the optimizer update, and the trainer waits while rollouts are produced. For simple environments that may be acceptable. For agentic RL, generation can involve:

  • code execution;
  • browser/search calls;
  • repository operations;
  • long reasoning traces;
  • multiple environment turns.

Rollout generation becomes extremely expensive. So asynchronous training decouples the two: actor/generator pool continuously produces trajectories; learner pool continuously consumes them and updates the policy. IBM’s Granite 4.2 implementation literally uses separate GPU pools and streams updated weights back to the rollout workers while they remain active. (Hugging Face) The utilization gain comes with policy staleness: a queued trajectory generated by θₜ₋₃ may reach a learner already optimizing θₜ. Faster policy changes and longer rollouts increase that gap. A larger rollout buffer can improve utilization while retaining more stale experience; a small buffer or strict synchronization keeps experience fresher but can leave compute idle.

A robust system therefore needs to make policy version a first-class property of every trajectory. Depending on the algorithm, you may:

  • discard very stale samples;
  • bound buffer age;
  • reweight off-policy samples;
  • clip updates more aggressively;
  • or deliberately slow policy drift.

This becomes especially important for long-horizon agents because one trajectory may take minutes while the learner performs several updates elsewhere. At architect level, the useful mental model is that RL infrastructure is a distributed system with consistency semantics. The distributed-systems analogy is useful for synchronization and version tracking, but database consistency guarantees do not describe RL convergence. Record the behavior-policy version and probabilities for each rollout, bound staleness, and validate the algorithm’s off-policy correction; higher utilization is valuable only if learning remains stable.


Worth Your Time

OpenAI. Jalapeño’s first measured results

Read the matched-latency/per-watt methodology and the architecture section, then pair it with:

SemiAnalysis. OpenAI Jalapeño technical analysis

The disagreement/caveats are as useful as the raw benchmark.

IBM. Granite 4.2 technical walkthrough

The sections on staged agentic RL and asynchronous GRPO are more valuable than the leaderboard tables.

Selective Regenerative Decoding

Read it for the partial-trajectory repair abstraction rather than the benchmark delta.

Google. Gemini Enterprise for Legal

Use it as a reference design for vertical enterprise agents: skills + permission-preserving connectors + agents + governance.

Today’s architectural takeaway: co-design is moving up and down the whole AI stack. Hardware is being designed around real agent traffic. Models are being post-trained inside actual execution environments. Vertical products are embedding permissions and domain procedures directly into the runtime. Even test-time reasoning is starting to preserve good state and repair only the failed portion.

The practical principle is:

When a layer becomes expensive enough, stop treating its neighbors as fixed. Optimize the boundary between them.