Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH RiVER Enables Reinforcement Learning Without Ground-Truth Labels
RESEARCH World Model Hallucination Is a Data Problem, Not Architecture
RESEARCH Single Researcher Places 2nd in ICRA Robot-Folding Challenge
RESEARCH Models Shed Learned Rules During Training
RESEARCH Free Scoring Signal Emerges from Standard RL Post-Training Runs
RESEARCH DeepMind Forensic Protocol Diagnoses Confused vs. Misaligned AI
RESEARCH Multimodal Models Flip Answers When Evidence Order Changes
RESEARCH Production Voice AIs Ignore Emotion, Approving Fraud and Ending Care Calls
RESEARCH Qwen's 397B Model Simulates Agent Environments Better Than GPT-5.4
RESEARCH FFASR Benchmark Exposes Far-Field Speech Recognition Gap RESEARCH Strict Regex Fix Raises Agent Grading Recall by 60 Percentage Points
RESEARCH InSight Enables Robots to Autonomously Learn New Tasks
RESEARCH OpenThoughts-Agent Dataset Hits 44.8% on Agentic Benchmarks
RESEARCH Amortized In-Context Learning Cuts Few-Shot Serving Cost
RESEARCH Google DeepMind's DiffusionGemma 28.6X harder to interpret than autoregressive models
RESEARCH Princeton Releases LOCUS, Machine-Readable Corpus of 9,239 US Local Ordinances
RESEARCH Three Agents Beat Every SQL Benchmark With Zero Fine-Tuning
RESEARCH OpenAnt LLM Pipeline Flags 28 Exploitable Vulnerabilities in OpenSSL RESEARCH MIT Extracts Attention Logic Into Swappable Python Code RESEARCH Physics-Augmented Koopman Networks Guarantee Generalization on Irregular Meshes
RESEARCH Single Dense Model Hosts Hundreds of Agent Personas as Lightweight Masks
RESEARCH Sparse Attention Heads Redirect Vision-Language Models With 83% Accuracy
RESEARCH ClinHallu Dissects Why Medical LLMs Misread Images 65% of the Time
RESEARCH Component Interaction, Not Quality, Determines Agent Performance
RESEARCH DiffusionGemma's Actual Decoding Contradicts Google's Block-Autoregressive Claims
RESEARCH DeepMind's Report Names "Jagged" Capability Gains as ASI Risk
RESEARCH Label-Free Test Catches LLM Reasoning Failures Better Than Self-Consistency
RESEARCH Agents-K1 Replaces RAG Text Chunks With Typed Scientific Knowledge Graphs
RESEARCH EvoArena Benchmark Exposes Agent Collapse in Evolving Environments
RESEARCH Sub-$11 Agent Outperforms Specialized Research Frameworks
RESEARCH Recursive Agent Harness Achieves 89% Accuracy on Long-Context Code Tasks
RESEARCH Half of AI-Generated Code Fixes Fail Human Review
RESEARCH DIRECT cuts embodied AI latency 65% with dynamic planner routing
RESEARCH Claude Fable 5 Autonomously Patched Code and Cost $110 in a Day
RESEARCH New Tool Finds 1,060 Hidden Training Dependencies Across Major LLMs