Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH
RESEARCH
RESEARCH LLM Agents Fall to Supply-Chain Attacks Hiding in Plain Language
RESEARCH Inference-Time Scaffolding Lifts Weak Model Accuracy to 0.91
RESEARCH IBM's VAKRA Benchmark Shows Models Fail 97% of Policy-Constrained Tasks RESEARCH RCI Framework Cuts Constraint Violations in Offline Safe RL Training
RESEARCH Liquid AI ships 3.1B edge vision model with 228-token/second throughput
RESEARCH Open Geospatial Embeddings Lower Barrier to Satellite Data Analysis
RESEARCH AI System Narrows Bounds on 40-Year Math Problem
RESEARCH Unified LLM Training Reveals Fundamental Mode Conflict
RESEARCH MMDiff Lets Engineers Audit and Edit Multimodal Model Features
RESEARCH ArchAgent v2 beats human champions on three-level cache design RESEARCH GENCO replaces three classical power grid solvers with one neural architecture
RESEARCH Amazon's Consilience Detects Silent Mode Collapse in Confidence-Based Reasoning
RESEARCH Taboo Stress-Tests LLMs on Production Constraints
RESEARCH Muon Optimizer Collapse After Grokking Threatens Production Training
RESEARCH SABRE Benchmark Exposes Vision Models' Blindness to Visual Contradictions
RESEARCH CreativeInstruct Recovers Diversity Without Slowing Inference
RESEARCH Mankind Labs Cuts Agentic Loop Tokens 17–26% With Reversible Memory Eviction RESEARCH CoinRAG Achieves 5.3% F1 Gain on Multi-Hop RAG with Nugget Caching
RESEARCH CalibForge Trains Models to 30-Point Gains with Solver-Calibrated Tasks
RESEARCH AV-AIVAT Cuts Agent Evaluation Cost 74-Fold With Certified Stopping RESEARCH Programmatic Tool Calling Beats JSON on New AI Models RESEARCH NVIDIA Publishes Full Greek RAG Playbook Reaching 2.3× Retrieval Gain
RESEARCH DeepMind's WeatherNext Cyclones Outpace Hurricane Models by Full Day
RESEARCH Skill Entropy Lifts Reasoning Model Scores From 34% to 68%
RESEARCH OctoLong improves code agents with cross-repository training
RESEARCH EvolveNet Evolves Agent Harness Instead of Model Weights
RESEARCH Chain-of-Thought Works at Tiny Scale, Breaking Compute Link RESEARCH SeGaBench Tests LLMs on Compiler-Blind Code Optimization
RESEARCH ReflectRL Recovers Reasoning Gains from Failed Expert Rollouts
RESEARCH Test-time scaling regimes require distinct evaluation profiles RESEARCH ALiBi Models Lose Positional Awareness in Long-Context Inference
RESEARCH Launch Inputs Matter More Than Model Choice for TUI Testing