LLM Semantic Router releases Decision 3.0: 5 open multimodal decision models, 0.8B to 27B, with d3 scoring 64.1.
Sakana AI’s Multi-Layered Review catches 73% of core-claim errors in a new 1,164-error benchmark for LLM-assisted peer review ...
Discover how autonomous AI agents escaped a secure sandbox and breached Hugging Face infrastructure without any human ...
OrcaRouter's OrcaCyber Zero 1.5 is a gated 1M-context cybersecurity model scoring 100% on Cybench, 95.8% on CVE-Bench subset.
OpenAI’s Decisions API enters public beta, returning typed answers about 10x faster at $0.10 per 1M input tokens.
Microsoft-Decision-1 is a Qwen3.5-9B decision-scoring model returning calibrated option probabilities at 85 ms p50 for $0.042 ...
OrcaRouter released OrcaCyber Zero 1.5, a gated cybersecurity model with a 1M-token context window. It reports 100% on Cybench and 95.8% on an evaluable CVE-Bench subset, priced at $3.00 / $7.50 per ...
Sakana AI’s TMLR paper introduces Multi-Layered Review, a 3-agent Claude-based reviewer, and a 1,164-error Contradiction Benchmark. MLR caught 73.43% of core-claim ...
JetBrains Mellum2.1 is an open 12B MoE coding agent model with 2.5B active parameters and 47.0% SWE-bench Verified.
Architect's Liquid Inference auctions every LLM request across competing providers, locking a max price before the first ...
NVIDIA's PivotOPD trains multi-turn agents to recover from pivotal mistakes and beats 13 baselines on ALFWorld, WebShop, ...
Google Cloud’s Gemini agent is one enterprise agent routing work across Gemini and Claude models, with spend caps.