Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH Agents-K1 Replaces RAG Text Chunks With Typed Scientific Knowledge Graphs
RESEARCH EvoArena Benchmark Exposes Agent Collapse in Evolving Environments
RESEARCH Sub-$11 Agent Outperforms Specialized Research Frameworks
RESEARCH Recursive Agent Harness Achieves 89% Accuracy on Long-Context Code Tasks
RESEARCH Half of AI-Generated Code Fixes Fail Human Review
RESEARCH DIRECT cuts embodied AI latency 65% with dynamic planner routing
RESEARCH Claude Fable 5 Autonomously Patched Code and Cost $110 in a Day
RESEARCH New Tool Finds 1,060 Hidden Training Dependencies Across Major LLMs
RESEARCH Token-Level Branching Offers Faster LLM Agent Training Without Budget Expansion
RESEARCH Token Recovery Closes Accuracy Gap While Halving VLM Inference Compute
RESEARCH Tahoe Text-to-SQL System Cuts Compiler Feedback by 96%
RESEARCH Google's DiffusionGemma Hits 1,000 Tokens Per Second
RESEARCH GRPO Cuts Pause-Handling Errors in Full-Duplex Agents Without Semantic Loss
RESEARCH Kamai's Phase Diagram Predicts Multimodal Failure Before GPU Commit
RESEARCH ABC-Bench Shows LLM Agents Now Outperform Expert Biologists on Lab Tasks
RESEARCH FPCG steers reasoning models at test time without retraining
RESEARCH Linear Probes Predict Reasoning-Model Behavior at 64–91% Accuracy RESEARCH LLM Leaderboards Fail to Predict Production Reliability
RESEARCH Grok 3 Surpasses Credentialed Biologists on Autonomous DNA Lab Tasks
RESEARCH EEVEE Surpasses Self-Improving Agents with 48% Margin on Multi-Domain Inference
RESEARCH Piper Compiler Eliminates Hand-Coding for Distributed Training
RESEARCH Single Linear Layer Outperforms 1M-Parameter Gate in MTP Speedup Test
RESEARCH Real EHR Benchmark Exposes Limits of LLMs in Clinical Action
RESEARCH AHA-WAM achieves 4.59× faster robot control by decoupling Diffusion Transformers
RESEARCH FASE Cuts Hallucination Detection to 333x Speed
RESEARCH New DRPO Method Fixes Long-Tail Vocabulary Collapse in LLM RL
RESEARCH FASE Cuts Hallucination Detection Cost to 0.3% of Rivals
RESEARCH SIGA Speeds Coding Agents on Scientific Simulators by 36×
RESEARCH Echo-Memory Shows World Models Fail the Revisit Test
RESEARCH Waterloo researchers cut uncertainty quantification cost 99.7% with FASE
RESEARCH EvalCards Schema Exposes Systematic AI Benchmark Metadata Gaps
RESEARCH Perplexity Agentic AI Cuts Task Time 87 Percent in Production Study
RESEARCH 64 Percent of Audio-Text Conflicts in AI Models Are Fixable
RESEARCH Router Matching 50 Retries with 10 Samples Cuts LLM Test-Time Compute
RESEARCH StreamMA Cuts Multi-Agent Reasoning Latency 26.9×
RESEARCH Vendor-Diverse Judge Panels Eliminate Bias in Language Model Evaluations