Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH
RESEARCH
RESEARCH Jev choice classification runs in 25 lines of Python with local inference
RESEARCH Claude Opus 5 leads circuit design benchmark at 61.6%, still fails 4 in 10 RESEARCH Cognition's SWE-2 matches frontier code performance at 64% lower cost
RESEARCH Knowledge Pull Requests ground QA 69% on revised articles versus 48% for raw rewrites
RESEARCH GravityOCR achieves 3.94× document OCR speedup via speculative decoding
RESEARCH GPT-6 Astra's compressed reasoning traces don't transfer to weaker models
RESEARCH Microsoft's Agensh reaches 55% test-pass rate with 1,024 agents
RESEARCH SWE-Serve benchmark reveals 23-point gap between local tests and production serving
RESEARCH A2M attack hijacks MCP agents with poisoned tool metadata at 93.6% success
RESEARCH Harness choice drives 5x cost gap while coding agent success rates stay flat
RESEARCH Flash-dLLM cuts diffusion LLM inference latency 11× with fused kernels
RESEARCH Six ASR Models Memorize Benchmarks Instead of Transcribing
RESEARCH Google DeepMind paper shows when routing signals waste compute
RESEARCH MidTool's 20B-Token Corpus Lifts Tool-Use Performance Across Benchmarks
RESEARCH VLA Detects Hidden Agent Coordination at 0.993 Accuracy
RESEARCH SPADE Agents Learn By Teaching Themselves New Tasks
RESEARCH Liquid AI's DSpark hits 3.18x H100 speedup with speculative decoding
RESEARCH New Method Proves Model Reasoning Path in 128 Test Cases, Fails on Standard Models
RESEARCH IBM Study: Curated Memory Boosts Mid-Tier Agents 16 Points RESEARCH METR Finds AI Accelerates Vulnerabilities But Not Optimization
RESEARCH BATON Cuts Robot Task Exploration Cost from Exponential to Linear
RESEARCH OpenAI and NVIDIA Lock 8-Gigawatt AI Factory to Ohio Campus
RESEARCH Memory Allocation Fix Lets Recurrent Models Match Attention at Scale
RESEARCH Researchers Expose Model Hypnosis: Invisible Prompt Cues Hijack AI With 99% Success
RESEARCH Compliance Detectors Fail Rule-Blindness Test, Lexsi Audit Finds
RESEARCH KV-Rescue Recovers 87% of Lost Model Accuracy
RESEARCH QuoteBench Exposes 55-Point Gap Hidden in Coding-Agent Benchmarks
RESEARCH Google DeepMind Ships Sign Language AI to Pixel 11 Users RESEARCH Agents Solve Only 27 of 43 Verified Code Tasks
RESEARCH LLM Agents Fall to Supply-Chain Attacks Hiding in Plain Language
RESEARCH Inference-Time Scaffolding Lifts Weak Model Accuracy to 0.91
RESEARCH IBM's VAKRA Benchmark Shows Models Fail 97% of Policy-Constrained Tasks RESEARCH RCI Framework Cuts Constraint Violations in Offline Safe RL Training
RESEARCH Liquid AI ships 3.1B edge vision model with 228-token/second throughput