Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH Every Guardrail Classifier Tested Fails Formal Safety Verification
RESEARCH Math Proof Shows Transformer Attention Stabilizes Predictably
RESEARCH AI Agents Bypass Software Engineering, Risk Production Failure
RESEARCH SLIM improves LLM agent performance 7 percentage points
RESEARCH Shepherd Raises Agent Accuracy 90% With Forking Traces
RESEARCH WildClawBench: Claude Opus Clears 62% in Real-World Agent Evaluation
RESEARCH Sparse MoE Models Match Dense Transformers at 3× Faster Inference
RESEARCH Muon Optimizer Achieves 2× Speed Over AdamW in Production LLM Training
RESEARCH CIVeX Logs Zero False Executions in Confounded Workflows
RESEARCH Paper Dismantles Causal Discovery Claim in Prediction Models
RESEARCH Frozen Models Encode Semantic Roles Without Fine-Tuning
RESEARCH Flow-OPD Raises Stable Diffusion Accuracy to 92 From 63
RESEARCH Conformal Path Reasoning cuts knowledge graph answer sets by 40 percent
RESEARCH Los Alamos Team Trains 8B Model That Generalizes Across Reasoning Benchmarks
RESEARCH Longer Context Degrades LLM Cooperation, Study Finds
RESEARCH AutoTTS Cuts Inference Costs 69.5% With Learned Test-Time Scaling
RESEARCH Coupling Tax: Reasoning Mode Cuts Accuracy Under Token Limits
RESEARCH Frontier Models Disagree on Ambiguous Policies, DRIP-R Shows
RESEARCH ActCam Controls Video Cameras and Characters Without Fine-Tuning
RESEARCH Optimizer-Model Consistency Cuts LLM Forgetting in Finetuning
RESEARCH Math AI Training Solver Accuracy Rises 21.4% With Verifier-Backed Generation
RESEARCH Arena Analysis: 66% of Leaderboard Votes Cancel Out
RESEARCH SIRA Outperforms Dense Retrieval Without Training or GPU Infrastructure
RESEARCH UniPool cuts MoE parameter budget 34 to 58 percent
RESEARCH DeepMind Math AI Hits 48% on Research-Grade Problems
RESEARCH Sparse MoEs retain accuracy at 87.5% weight pruning
RESEARCH Rice and Apple researchers cut image-generation FID 22% with token fix
RESEARCH First-Token Entropy Rivals Multi-Sample Hallucination Detection
RESEARCH MRI-Eval Finds LLMs Score 97% on Flashcards, 30% on Open Recall
RESEARCH Q2RL Reaches 100% Success on Peg Insertion, Outpacing BC and IBRL
RESEARCH Purdue and Georgia Tech Prove Transformers Extract Nonlinear Features in Context
RESEARCH LongSeeker Beats Competitors on Long-Horizon Tasks
RESEARCH Verkor's Agentic System Closes RTL-to-Layout in 80 Hours
RESEARCH Multi-Agent LLMs Lose One-Third Quality But Signal Recovery Path
RESEARCH Claude's Safety Tests Fail When Model Hides Suspicions Inside
RESEARCH Mozilla Found 12 Critical Firefox Bugs Using Claude Mythos AI
RESEARCH SymptomAI Outperforms Clinicians 2.47x in Real-World Trial
RESEARCH OpenSeeker-v2 beats Alibaba's Tongyi on agentic search benchmarks
RESEARCH Automated agent recommender cuts multi-agent system engineering steps to one
RESEARCH Dreadnode Framework Cuts AI Red Teaming from Weeks to Hours