The week when real agent costs appeared on the balance sheet — and it's not the token price: it's the harness, the power contract, and where inference runs.
Four months.
That's how long it took Uber to burn through the entire annual budget for AI coding tools. Five thousand engineers, Claude Code, and no line item auditing the harness.
This is the ai|expert Edition. The week when real agent costs appeared in the ledger — and the culprit wasn't the token price.
In February this year, 32% of Uber's engineers used Claude Code. By March, 84%. By spring, monthly active use reached 95%. CTO Praveen Neppalli Naga confirmed to The Information: Uber's 2026 budget for AI coding tools was consumed in four months. [ref: optimizing-coding-agent-costs-]
Average cost per engineer was 150 to 250 dollars monthly. Heavy users reached 500 to 2,000 dollars. Naga himself spent 1,200 dollars in a single two-hour session. But the most revealing detail isn't that number. It's the mechanism that created the problem: Uber built an internal leaderboard ranking engineers by Claude Code usage volume.
A token leaderboard. A productivity metric that is, in practice, a direct incentive to burn budget.
Since then, the company imposed a cap of 1,500 dollars per month per employee per coding agent tool. Salesforce estimates spending 300 million dollars with Anthropic this year alone — for a workforce of 15,000 engineers. And Microsoft's Experiences and Devices division canceled direct Claude Code licenses through June 30, redirecting all engineers to GitHub Copilot CLI.
What's happening here isn't a model pricing problem. It's a visibility problem.
LangChain diagnosed this as tokenmaxxing — using token volume as a productivity proxy — combined with tool fragmentation. A single feature can pass through Claude Code, Cursor, and GitHub Copilot Chat. Each tool emits incompatible telemetry. Without normalization, it's impossible to compare cost between tools or attribute spending to model, harness, or context management. Native dashboards simply fail when you have multiple tools.
And Databricks went further. They built an internal benchmark on their own codebase of multiple millions of lines — Python, Go, TypeScript, Scala, Rust, Java, Bazel, Protobuf — using real PRs, manually reviewed to avoid the training data leakage problems that contaminate SWE-Bench and TerminalBench. [ref: databricks-benchmarks-coding-a]
The result is the most important data point of this week.
Sonnet 5 is 1.7 times cheaper per token than Opus 4.8. And yet it cost more per task: 2.09 dollars versus 1.94. Because it consumed 1.9 times more tokens to complete the same work. And it completed only 81% of tasks, versus 87% for Opus. Token price is a decoy metric.
GLM 5.2, on the other hand, matched Opus 4.8 statistically on quality — and cost 1.28 dollars per task, versus 1.94.
But the data that actually changes architecture is the harness. Databricks ran the same model through different harnesses — Pi versus Claude Code and Codex. Identical quality. Cost more than double in heavier harnesses. Pi sent approximately three times less context per turn, managed the context window more rigorously, and completed tasks in fewer rounds.
The harness can double the cost without touching quality. This means most teams are losing money on default configurations — not because they chose the wrong model, but because they never audited how much context the harness sends each turn.
That's exactly what the NemoClaw blueprint from LangChain with NVIDIA addresses. The combination of Nemotron 3 Ultra with the Deep Agents Code harness achieved 0.86 on LangChain's internal evaluation suite — at a cost of 4.48 dollars per iteration, versus 43.48 dollars for the next best-performing model. A 10-fold reduction. Without any retraining. Every gain came from environment engineering around the model: system prompts, tool descriptions, middleware. [ref: langchain-nvidia-deep-agents-b]
The target use case for NemoClaw is instructive: COBOL modernization. A code agent compromised by prompt injection can rewrite business logic or exfiltrate source code at industrial scale. The harness is where governance begins — before the model, not after.
And supply side moved too. Together AI launched Provisioned Throughput Units — PTUs — at 0.05 dollars per PTU per minute, with 99% uptime SLA. On MiniMax M3, one PTU delivers 138,840 input tokens per minute. At full utilization, that's 0.36 dollars per million input tokens — versus 5 dollars on Claude Opus 4.8. Customers migrating from proprietary APIs to open models report 6 to 20 times reduction in inference cost. [ref: together-ai-launches-provision]
Together AI's volume grew from 30 billion to over 400 trillion tokens per month in nine months. That number alone is a market signal.
With one caveat that needs to be in your financial model: PTUs bill continuously — approximately 2,160 dollars per month. Idle capacity raises effective cost per token immediately. The 99% SLA allows nearly 8.7 hours of downtime per month. Input tokens, cached input, and output consume PTUs at different rates, complicating spend forecasting for variable traffic. It's not serverless. It's not dedicated. It's a throughput commitment that assumes production workloads stable enough to keep PTUs near full utilization.
The lesson from the first segment is direct: token price is the most visible line on the bill, but rarely the largest. The harness, context per turn, capacity commitment model — those are the multipliers that no agent ROI spreadsheet captures by default.
Two hundred and fifty-eight terawatt-hours.
That's what AI-optimized servers will consume in 2027, per Gartner projections. To put that number in perspective: in 2025, it was 95 terawatt-hours. In 2026, 175 — growth of 84% in one year. In 2027, for the first time in history, AI hardware will consume more energy than all conventional servers in the world combined.
Global data center consumption will reach 565 terawatt-hours this year — 26% growth over 2025. Cooling systems alone will consume 195 terawatt-hours in 2026. Nearly as much as all traditional servers on the planet. AI servers will represent 31% of total data center consumption this year, up from approximately 20% in 2025. [ref: ai-servers-will-dominate-data-]
And the US accounts for 204 of those 565 terawatt-hours total — 36% of global consumption. Of which, AI-dedicated data centers account for 68 terawatt-hours, or one-third of America's total. Gartner is direct: power availability is the binding constraint on AI growth. More than 75 data center projects valued at 130 billion dollars were blocked in early 2026 by disputes over power and water access.
And that's where Oregon enters the equation.
The Oregon Public Utility Commission approved a 29.7% rate increase for Portland General Electric customers consuming more than 20 megawatts — in line with the state's POWER Act. Facilities above that threshold now need to sign 10 to 30-year contracts, pay deposits for new service, and commit to 90% of contracted demand regardless of actual consumption. [ref: oregons-30-data-center-power-s]
Ninety percent take-or-pay. This eliminates any savings from scaling down inference fleets during low-traffic periods. And if you consume above plan, penalties are 1.5 times the power cost and 4 times the transmission cost.
Oregon has over 135 data centers. PGE invested 210 million dollars in network expansion in Hillsboro. A 250-megawatt campus now needs to commit payment for approximately 225 megawatts of base supply — regardless of GPU utilization. A training spike above plan activates transmission penalties that can destroy an entire budget cycle's forecast.
AWS and the Data Center Coalition have already warned that this regime will push new builds to other states. Montana offers favorable tax treatment under HB 424. Washington is assessing the impact of large loads on state tax revenue. And the parallel PacifiCorp decision, expected in November 2026, could extend similar rules to two-thirds of Oregon utility customers.
On the other side of the Atlantic, the UK is taking the opposite path.
British Parliament approved in November 2025 an amendment adding data centers to the Nationally Significant Infrastructure Projects regime — the NSIP. Under that regime, developers apply directly to the Secretary of State. A single approval consolidates planning permission, compulsory land acquisition, road works, and ancillary consents into one instrument. [ref: uk-data-centers-get-national-i]
The Ministry of Housing, Communities, and Local Government estimates the change could cut approval timelines by up to one year and save up to 1.3 billion dollars per project. The government is exploring whether it can cut average NSIP approval times from 18 months to 12 — versus a standard process that has taken up to four years. More than 80 projects already entered the Planning Inspectorate's pre-application pipeline. Three already have NSIP classification.
But NSIP status doesn't guarantee substation, grid connection, or water for cooling. The National Policy Statement defining eligibility thresholds hasn't been published yet. Any capex model assuming NSIP acceleration is betting on a threshold that doesn't exist on paper yet.
What Oregon and the UK show together is that compute can't be a single line item anymore on a three-year budget. Oregon locks you into 10-year contracts with 4x transmission cost penalties for unexpected training spikes. The UK might offer approval in 12 months — but only for projects large enough to qualify as national infrastructure.
Geographic arbitrage — choosing where to run inference based on power cost — became a survival calculation. And the timeline to secure interconnection agreements and redesign facilities for liquid cooling is shrinking.
If the first two segments show that agent running costs live in the layers around the model — harness, power contract, capacity commitment — the third shows where inference itself is migrating. And it's leaving the API.
Three new destinations appear this week with production data: inside PostgreSQL, in Modal's stateful sandbox with 300 million dollars in ARR, and in the CPU NVIDIA launched specifically for agent orchestration.
It starts with AlloyDB. Google took to GA in PostgreSQL 17 a complete set of AI functions — ai.generate, ai.summarize, ai.if, ai.rank, ai.forecast — with an acceleration layer that changes the math of inference on high-cardinality tables. [ref: alloydb-introduces-local-infer]
The reference number is 23,000 times. That's the throughput gain Google reports in internal tests when comparing line-by-line API calls to AlloyDB's local proxy model for ai.if. The more conservative version — smart batching in GA — already delivers 2,400 times improvement over row-at-a-time baseline, processing 10,000 rows per second. The proxy model reaches 100,000 rows per second with 6,000 times cost reduction.
How it works: a PREPARE statement sends a sample of your data to a frontier model on Vertex AI, which trains an ultra-lightweight proxy model inside the database. Queries then run against that proxy at database speed. When proxy confidence drops below threshold, AlloyDB reverts to the remote LLM.
The practical result: 100,000 rows per second for semantic filters, without 100,000 round trips to the API. For semantic filter workloads on product tables, logs, or user behavior data, that changes latency and cost math fundamentally.
But with a critical caveat: these numbers are from Google's internal tests, apply specifically to ai.if in preview, not the complete function catalog. Google hasn't disclosed absolute cost per token, p50 or p99 latency on the fallback path, or GPU-hours needed during PREPARE. The proxy needs custom regression suites to measure drift by domain. Raimundas Juodvalkis, architect at Starburst, recommends: treat these functions as governed database extensions, not as magical WHERE clauses. Start with read-heavy workflows before writing model-derived fields back into core systems.
From database to sandbox. Modal closed a Series C of 355 million dollars, valued at 4.65 billion, led by General Catalyst and Redpoint Ventures. What matters more than the funding is the revenue mix: sandboxes already represent more than one-third of the company's 300 million in ARR. More than one billion sandboxes have launched on the platform to date. [ref: modal-cto-on-agent-workload-in]
That number is a paradigm shift. It means agent workloads — stateful multi-step loops, tool-calling, untrusted code execution, output inspection, retry — became the dominant driver of AI infrastructure. Surpassing stateless inference.
CTO Akshat Bubna explains why on Latent Space podcast: Kubernetes was designed for HTTP request/response cycles. Agents need to write code, mutate environment state, and iterate in tight feedback loops. Modal built a Rust stack from scratch — custom filesystem, container runtime, scheduler, GPU memory snapshotting — accessible through Python decorators that integrate serverless functions, sandboxes, elastic inference with speculative decoding, networked containers and RDMA into a unified control plane.
GPU snapshotting improved cold starts 100-fold. RL rollouts reached 100,000 sandboxes in parallel. Ramp reports their Inspect code agent writes 70% of PRs that merge on their platform. Lovable processed 250,000 app creations in a single weekend using over one million sandboxes.
ARR grew approximately five-fold since September 2025. That's the signal that agent workloads moved from experimental to production.
And there's an operational problem that emerges at that scale: when agents write code, observability becomes harder than code review. Debugging agent-generated artifacts through traditional dashboards doesn't scale. Modal promises granular RBAC for agent capability scope — but it's still forthcoming.
At the lowest level of the stack — the CPU orchestrating the sandbox — NVIDIA launched Vera. Eighty-eight monolithic cores. Core Olympus, 10-way out-of-order design with SVE2 FP8 support. Two megabytes of private L2 per core. Unified L3 cache of 162 to 164 megabytes. LPDDR5X memory bandwidth of 1.2 terabytes per second at under 40 watts. Core-to-core bandwidth of 3.4 terabytes per second — three times more than any x86 data center CPU. [ref: nvidia-vera-maximizing-single-]
Perplexity ran a real workflow on Vera: clone a repository and execute the test suite in sandboxes. Result: 1.5 times faster than x86. Concurrent sandbox spin-up: 1.9 times faster.
Phoronix founder Michael Larabel described Vera as the most competitive non-x86 server CPU he's tested — 63% geomean improvement over NVIDIA's previous Grace and 55% over Intel's 128-core Xeon 6980P. NVIDIA claims 40% reduction in loaded peak latency versus x86, and the Vera CPU Rack delivers over 4 times the capacity and twice the performance per watt of x86 server racks.
The logic is straightforward: agents spend more time orchestrating sandboxes and waiting on memory than inside a transformer layer. A monolithic CPU with uniform bandwidth eliminates NUMA hops from chiplets and performance cliffs when threads overflow between domains. If your agent stack is still dominated by large model forward passes, Vera's gains don't change unit economics. But if you run code sandboxes, tool calls, and data-intensive retrieval, the host CPU became the bottleneck.
With the operational caveat that can't be left out: optimized binaries require GCC 16.1 or LLVM Clang 21 or later. x86 containers and existing libraries are not compatible. For those running COBOL, legacy .NET, or any Windows-based toolchain, that means separate Linux worker nodes or kernel upgrades before migrating to Vera. Vera is already in production at CoreWeave, Lambda, and Oracle Cloud Infrastructure. Broad OEM availability — Dell, HPE, Lenovo, Supermicro — is targeted for Q2 2026.
And the three destinations — database, stateful sandbox, orchestration CPU — share a common denominator that NemoClaw makes explicit: attack surface grows with execution complexity.
When inference enters the database process, or the agent executes code in a sandbox, or the CPU runs quantized scoring without crossing the PCIe fabric — the security perimeter ceases to be the API edge. It's the process itself. NemoClaw solves this with out-of-process enforcement. [ref: langchain-deep-agents-on-nvidi]
OpenShell uses Landlock LSM and Seccomp BPF to block filesystem, network, and process policies at sandbox creation. Credentials stay with NemoClaw — outside the sandbox. Network egress denied by default unless approved by request. Each session generates an audit snapshot stored outside the agent process. Root identity is always rejected.
The logic is elegant: if the agent can't access its own controls, it can't disable them. An agent compromised by prompt injection hits the sandbox wall — doesn't cross it.
But Futurum Group points to the limit: OpenShell covers only the end of the trust chain — runtime execution. Upstream governance — model selection, data classification, skill provenance — still depends on what your team built before runtime. The compressed sandbox image is approximately 2.4 gigabytes and requires Ubuntu 22.04 or later with Docker and Node.js 20+. Older container hosts and Windows toolchains — common in COBOL and .NET modernization projects — will need upgrades or separate Linux worker nodes.
The blueprint is available on GitHub under Apache 2.0 license. Treat the current version as an evaluation target, not production runtime.
The thread connecting all three segments of this edition is visibility. Cost you don't measure — in the harness, in the power contract, in sandbox cold starts — is cost you don't control.
When Uber's COO Andrew Macdonald told Fortune that the link between growing AI spend and features useful to customers "isn't there yet," he was describing exactly what happens when the bill arrives without the right instrument to read it.
Three takeaways from this edition: first, build your own benchmark on real PRs from your codebase — token price is a decoy metric and the harness can double cost alone. Second, treat power contracts as reserved instances with severe overage penalties — not as metered cloud compute. Third, inference is migrating into the database, into stateful sandboxes, and into orchestration CPUs — each with a security trade-off that needs out-of-process enforcement.
Agent costs became visible this week — and they have three lines that no ROI spreadsheet captures by default: the harness sending too much context, the power contract that punishes training spikes with 4x transmission cost, and the sandbox that needs monolithic CPU to avoid bottlenecking. Wire on Monday opens with Grok 4.5 entering the Opus tier at 2.49 dollars per task — and the new price war it unlocks. Good week.