Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH Soft-Prefix Attacks Flip LLM Reasoning at 90% on Hidden Vector Injection
RESEARCH Dense Patch Tokens Match Vision-Language Models at 1% the Parameters
RESEARCH SWE-Pruner Pro Cuts Coding-Agent Token Use 39%
RESEARCH OpenAI Suspends Model After Escaping Sandbox, Bypassing Security
RESEARCH Android Agent Framework Cuts Mobile Task Time by 95 Percent
RESEARCH E3 Method Cuts LLM Agent Token Use by 91% on Code Edits
RESEARCH TerraZero Tops InterPlan Benchmark Without Human Data
RESEARCH LLM Judges Reverse 85% of Verdicts When Given Reference Answers
RESEARCH Three Hours on $329 GPU Replaces Thousands of Hours of NAS Training
RESEARCH Apple's MM-ToolSandBox Reveals Why Half of Frontier AI Agents Fail on Visual Tasks
RESEARCH Activation-Level Fixes Outperform Prompt Edits for Biased LLM Judges
RESEARCH ZoRRO Matches Deep Learning CTR at 600× Speed
RESEARCH Super Weights Training Fails on OLMo Models, Demolishing Sparse Fine-Tuning Strategy
RESEARCH Training-Efficient Low-Rank Compression Sidesteps Serving-Speed Proof
RESEARCH Hugging Face Cuts Inference Attention Overhead 20-40% With Fused Kernels
RESEARCH Claude Opus Fails Half of Real-World Tasks in UniClawBench
RESEARCH Cornell's Co-LMLM Matches GPT-4o-Mini by Storing Facts in a Database RESEARCH Timestep Weighting Cuts Reward-Model Query Costs for Diffusion RLHF
RESEARCH STRACE Framework Boosts Multi-Agent Verification by 16 Points
RESEARCH DynaKRAG Boosts Multi-Hop QA Accuracy by Up to 5.78 Points
RESEARCH OpenAI Reveals 30% of SWE-Bench Pro Tasks Are Broken
RESEARCH DepthWeave-KV cuts LLM cache memory by 8.3x without retraining
RESEARCH SovereignPA-Bench Measures Whether AI Agents Protect User Boundaries
RESEARCH Dual Planning Loop Solves Semantic Gap in Hierarchical Robots
RESEARCH SearchGen-20K Teaches Visual Generators When to Search
RESEARCH CompactionRL boosts GLM coding agents 5–7 points on benchmarks
RESEARCH New Verification Method Hits 86.5% on Terminal-Bench Without Fine-Tuning
RESEARCH Simple Threshold Monitor Matches Complex LLM Safeguards in ICML Paper
RESEARCH Misaligned Coding Agents Evade Monitors in 93% of Gradual Attacks
RESEARCH LACUNA Shows Unlearning Methods Fail to Erase PII from Models
RESEARCH Language Labels Beat Scalars in Offline Robot Learning
RESEARCH Theoria bridges formal proof and LLM judges with auditable verification
RESEARCH Three major benchmarks inflate coding-agent scores, audit finds
RESEARCH AutoMem Training Doubles Agent Performance on Long-Horizon Tasks
RESEARCH One Layer Matches Full RL Post-Training on Qwen Models