Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH Simple Prompting Baselines Outperform Complex Supervision Methods
RESEARCH Researchers Close Gap Between AI Agents and Hand-Curated Skills
RESEARCH New Training Technique Improves LLM Confidence Calibration by 63%
RESEARCH Google Releases Zero-Shot Tabular Model but Hides Benchmark Data
RESEARCH Vision-language models route knowledge through just 2.5% of network RESEARCH AI Agents Double Repository-Level Merge Friction
RESEARCH Original-Language Context Recovers Accuracy Lost in Multilingual Cascades RESEARCH ENS Hits 10× Accuracy on Tough PDE Benchmarks Without Correction Loops
RESEARCH Mechanism Taxonomy Lifts LLM Moderation F1 by 5.4%
RESEARCH Open-Weight Pipeline Achieves 68% Accuracy Extracting Political Networks from News
RESEARCH Sequence Probability Fails as Production Inference Signal
RESEARCH RiVER Enables Reinforcement Learning Without Ground-Truth Labels
RESEARCH World Model Hallucination Is a Data Problem, Not Architecture
RESEARCH Single Researcher Places 2nd in ICRA Robot-Folding Challenge
RESEARCH Models Shed Learned Rules During Training
RESEARCH Free Scoring Signal Emerges from Standard RL Post-Training Runs
RESEARCH DeepMind Forensic Protocol Diagnoses Confused vs. Misaligned AI
RESEARCH Multimodal Models Flip Answers When Evidence Order Changes
RESEARCH Production Voice AIs Ignore Emotion, Approving Fraud and Ending Care Calls
RESEARCH Qwen's 397B Model Simulates Agent Environments Better Than GPT-5.4
RESEARCH FFASR Benchmark Exposes Far-Field Speech Recognition Gap RESEARCH Strict Regex Fix Raises Agent Grading Recall by 60 Percentage Points
RESEARCH InSight Enables Robots to Autonomously Learn New Tasks
RESEARCH OpenThoughts-Agent Dataset Hits 44.8% on Agentic Benchmarks
RESEARCH Amortized In-Context Learning Cuts Few-Shot Serving Cost
RESEARCH Google DeepMind's DiffusionGemma 28.6X harder to interpret than autoregressive models
RESEARCH Princeton Releases LOCUS, Machine-Readable Corpus of 9,239 US Local Ordinances
RESEARCH Three Agents Beat Every SQL Benchmark With Zero Fine-Tuning
RESEARCH OpenAnt LLM Pipeline Flags 28 Exploitable Vulnerabilities in OpenSSL RESEARCH MIT Extracts Attention Logic Into Swappable Python Code RESEARCH Physics-Augmented Koopman Networks Guarantee Generalization on Irregular Meshes
RESEARCH Single Dense Model Hosts Hundreds of Agent Personas as Lightweight Masks
RESEARCH Sparse Attention Heads Redirect Vision-Language Models With 83% Accuracy
RESEARCH ClinHallu Dissects Why Medical LLMs Misread Images 65% of the Time
RESEARCH Component Interaction, Not Quality, Determines Agent Performance
RESEARCH DiffusionGemma's Actual Decoding Contradicts Google's Block-Autoregressive Claims