aiexpert
Inicio / Podcast / Ep. 29
29
Episodio 29 · 4 min · Boletín

Kernel profiling, GPU safety, and agent harnesses became measurable this week—and the numbers expose where optimization actually lives.

Kernel profiling, GPU safety, and agent harnesses became measurable this week—and the numbers expose where optimization actually lives.

Presentan AlanPresentación
RSS

Transcripción del episodio

El guion que salió al aire, íntegro
Alan

Thirty percent.

Ada

That's how much a single matmul kernel shrinks when you can see where it stalls on memory. Google just gave you the tool. The question is whether your team knows how to read it.

Alan

This is the ai|expert Wire. Kernel profiling, GPU safety, and agent harnesses became measurable this week—and the numbers expose where optimization actually lives.

Alan

Google exposed cycle-level kernel profiling in XProf for Pallas this week. [ref: google-adds-cycle-level-kernel-profiling-to-xprof] Custom kernels on TPU v7 now surface hardware counters at microsecond resolution, letting teams see where compute stalls on memory and optimize accordingly. The profiler requires explicit flags and counter selection.

Ada

A Google demo cut a matmul kernel by 30%. [ref: google-adds-cycle-level-kernel-profiling-to-xprof] That's not a model improvement. That's visibility into where the actual work happens.

Alan

Nvidia opened native GPU programming in Rust this week with two different tracks. [ref: nvidia-announces-native-gpu-programming-in-rust-with-two-kernel-writing-tracks] cuda-oxide follows the familiar SIMT model with compile-time memory safety through DisjointSlice. cutile-rs uses a higher-level Tile model that eliminates thread races by construction. cutile-rs is stable on crates.io; cuda-oxide remains early alpha on nightly.

Ada

Memory safety by construction means you catch the race before it ships. That's the harness problem moving down the stack.

Alan

Data centers are shifting to 800-V distribution and deploying GaN semiconductors across multiple conversion stages to handle GPU currents approaching 10,000 A at sub-1 V core voltages. [ref: ai-power-demands-push-gan-into-data-center-design] GaN switching frequencies now reach 3 MHz—roughly 3× silicon's capability—with roadmaps targeting 10 MHz. [ref: ai-power-demands-push-gan-into-data-center-design]

Ada

Five kilowatt GPUs are arriving in two years. The power distribution that works today does not scale to that.

Alan

OpenAI's breach this week exposed the gap between model quality and operational control. A chained exploit of an unpatched image-processing vulnerability and SSO misconfiguration gave access to employee accounts and internal repositories. [ref: heap-overflow-and-sso-misconfiguration-compromised-openai-internal-repos] The same flaw affects Slack, Meta, GitHub Enterprise, and major frameworks.

Ada

But the real story is what the breach revealed about what matters in production. Five operational patterns—state ownership, mutation serialization, execution scoping, work bounding, and edge validation—replace model quality as the measure of production readiness. [ref: the-agent-harness-control-planes-invariants-and-approval-boundaries-for-producti]

Alan

Silent failures that users see but operators cannot explain are the failure mode that matters. [ref: the-agent-harness-control-planes-invariants-and-approval-boundaries-for-producti]

Ada

OpenAI's Agents API manages sessions, orchestration, and recovery for multi-step autonomous systems. [ref: openai-agents-api] But teams with strict data residency or retention requirements outside the US cannot use it, even with self-hosted sandboxes.

Alan

Three benchmarks this week measured what actually breaks in production. Claude Opus 5 leads circuit design at 61.6%, still failing 4 in 10 on real SPICE validation. [ref: can-ai-design-circuit-boards-yet] EEBench uses a declarative circuit language instead of GUI automation, grading against SPICE simulation and component tolerances. Frontier models can handle real electronics tasks but with substantial gaps.

Ada

Code agents show a 23-point gap between local tests and production serving. [ref: swe-serve-benchmarking-agentic-engineering-for-production-inference-serving] A third of patches pass all local tests but fail end-to-end validation on SGLang serving tasks. The benchmark spans 53 production-grounded tasks across six inference engineering families, exposing how multi-domain and concurrent coordination requirements expose agent failures that isolated testing misses.

Alan

Harness choice drives a 5× cost spread while success rates stay flat. [ref: harnesstax-how-much-does-the-harness-matter-for-coding-agents] A UC Berkeley and Arena Intelligence study of 21 model–harness pairs finds that Claude Code costs twice as much as Pi on the same models with nearly identical success rates.

Ada

The model matters less than the frame around it. That is the week's real number.

Alan

Visibility into where compute stalls, safety by construction in the kernel, and harnesses that measure what actually matters in production—that is infrastructure hardening. The Edition on Friday covers the full week of agent infrastructure, from MCP supply-chain risk to the scaling limits of decentralized coordination. Good week.