Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH Tahoe Text-to-SQL System Cuts Compiler Feedback by 96%
RESEARCH Google's DiffusionGemma Hits 1,000 Tokens Per Second
RESEARCH GRPO Cuts Pause-Handling Errors in Full-Duplex Agents Without Semantic Loss
RESEARCH Kamai's Phase Diagram Predicts Multimodal Failure Before GPU Commit
RESEARCH ABC-Bench Shows LLM Agents Now Outperform Expert Biologists on Lab Tasks
RESEARCH FPCG steers reasoning models at test time without retraining
RESEARCH Linear Probes Predict Reasoning-Model Behavior at 64–91% Accuracy RESEARCH LLM Leaderboards Fail to Predict Production Reliability
RESEARCH Grok 3 Surpasses Credentialed Biologists on Autonomous DNA Lab Tasks
RESEARCH EEVEE Surpasses Self-Improving Agents with 48% Margin on Multi-Domain Inference
RESEARCH Piper Compiler Eliminates Hand-Coding for Distributed Training
RESEARCH Single Linear Layer Outperforms 1M-Parameter Gate in MTP Speedup Test
RESEARCH Real EHR Benchmark Exposes Limits of LLMs in Clinical Action
RESEARCH AHA-WAM achieves 4.59× faster robot control by decoupling Diffusion Transformers
RESEARCH FASE Cuts Hallucination Detection to 333x Speed
RESEARCH New DRPO Method Fixes Long-Tail Vocabulary Collapse in LLM RL
RESEARCH FASE Cuts Hallucination Detection Cost to 0.3% of Rivals
RESEARCH SIGA Speeds Coding Agents on Scientific Simulators by 36×
RESEARCH Echo-Memory Shows World Models Fail the Revisit Test
RESEARCH Waterloo researchers cut uncertainty quantification cost 99.7% with FASE
RESEARCH EvalCards Schema Exposes Systematic AI Benchmark Metadata Gaps
RESEARCH Perplexity Agentic AI Cuts Task Time 87 Percent in Production Study
RESEARCH 64 Percent of Audio-Text Conflicts in AI Models Are Fixable
RESEARCH Router Matching 50 Retries with 10 Samples Cuts LLM Test-Time Compute
RESEARCH StreamMA Cuts Multi-Agent Reasoning Latency 26.9×
RESEARCH Vendor-Diverse Judge Panels Eliminate Bias in Language Model Evaluations
RESEARCH LLMs Can Induce Hidden Rules, but Procedural Execution Remains Uncracked
RESEARCH AdaCodec cuts video-token load by 7× with predictive encoding
RESEARCH SafeSteer cuts alignment tax by targeting sparse safety tokens
RESEARCH Output Format Drives Faster Accuracy Loss Than Domain Shift in Multimodal LLMs
RESEARCH SubFit Maintains 84.6% Accuracy While Pruning LLM Layers at 25% Sparsity
RESEARCH Claude Code Spent 58% of Sessions Optimizing a Broken Architecture
RESEARCH Robot Manipulation Accuracy Jumps 22.5% With Motion-Aware Encoder
RESEARCH Linear Inverse Problems Don't Protect Against Diffusion Hallucination
RESEARCH HullFT Method Cuts Test-Time Finetuning Latency Versus SIFT
RESEARCH GPIC Open-Source Dataset Displaces ImageNet-1K as Standard Training Corpus