Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH Allen AI's OlmoEarth v1.1 cuts satellite inference compute 3x
RESEARCH Medical LLMs Underweight Patient Autonomy
RESEARCH Researchers Map Hallucination Rates by Model Size and Data Frequency
RESEARCH RRFP Achieves 2.77× Throughput on Multimodal Pipeline-Parallel Training
RESEARCH DashAttention reaches 75% sparsity while matching full-attention accuracy
RESEARCH EnvFactory lifts Qwen3 tool-calling accuracy 15% with synthetic data
RESEARCH Memory Lookup Replaces Linear Attention Over Long Prefixes
RESEARCH SAEBench Metrics Rank SAEs Backwards, Audit Finds
RESEARCH Autonomous Disease Forecasting System Outperforms CDC Ensemble on Blinded Tests
RESEARCH FORGE Reduces Agent Failures to 1% Without Model Fine-Tuning
RESEARCH Grep Beats Vector Search in Inline Agent Retrieval
RESEARCH Frontier Agents Reach 25% on Real-World Forecasting Test
RESEARCH Microsoft Finds GPT-5 Fails Against Implausible Attacks
RESEARCH Scientific ML Models Disagree on 16% of Predictions Despite Matching Accuracy
RESEARCH LLM Formalization Catches 18.8% Ambiguous Requirements in Safety Specs
RESEARCH TFlow cuts multi-agent inference tokens 83% via weight injection
RESEARCH Negation Neglect Drives False Belief Rate to 88.6% in Fine-Tuned LLMs
RESEARCH Why Production Agents Fail Without Harness Infrastructure
RESEARCH Berkeley Framework Cuts Agent Latency 1.3–2.2×
RESEARCH KV-Fold Extends Transformer Context to 128K Without Retraining
RESEARCH IBM Boosts Zero-Shot Search Accuracy 25% With LLM Query Refinement
RESEARCH 27M Attractor Model Beats GPT o3 on Logic Puzzles
RESEARCH Reward Hacking Undetected in Single-Verifier Training
RESEARCH Sparse-to-Dense RL Lifts MATH Scores to 78.5% on Small Models
RESEARCH Standard load-balancing losses degrade SMoE expert specialization by 3x
RESEARCH VECA Cuts Vision Transformer Inference Cost to Linear Time
RESEARCH MEME benchmark finds 97% failure on agent memory dependency tasks
RESEARCH RuDE Predicts Fine-Tuning Success Without Training
RESEARCH Google's RubricEM trains research agents without ground truth
RESEARCH Every Guardrail Classifier Tested Fails Formal Safety Verification
RESEARCH Math Proof Shows Transformer Attention Stabilizes Predictably
RESEARCH AI Agents Bypass Software Engineering, Risk Production Failure
RESEARCH SLIM improves LLM agent performance 7 percentage points
RESEARCH Shepherd Raises Agent Accuracy 90% With Forking Traces
RESEARCH WildClawBench: Claude Opus Clears 62% in Real-World Agent Evaluation
RESEARCH Sparse MoE Models Match Dense Transformers at 3× Faster Inference
RESEARCH Muon Optimizer Achieves 2× Speed Over AdamW in Production LLM Training
RESEARCH CIVeX Logs Zero False Executions in Confounded Workflows
RESEARCH Paper Dismantles Causal Discovery Claim in Prediction Models
RESEARCH Frozen Models Encode Semantic Roles Without Fine-Tuning