OrcaRouter's OrcaCyber Zero 1.5 is a gated 1M-context cybersecurity model scoring 100% on Cybench, 95.8% on CVE-Bench subset.
Sakana AI’s Multi-Layered Review catches 73% of core-claim errors in a new 1,164-error benchmark for LLM-assisted peer review ...
Discover how autonomous AI agents escaped a secure sandbox and breached Hugging Face infrastructure without any human ...
Microsoft-Decision-1 is a Qwen3.5-9B decision-scoring model returning calibrated option probabilities at 85 ms p50 for $0.042 ...
OpenAI’s Decisions API enters public beta, returning typed answers about 10x faster at $0.10 per 1M input tokens.
JetBrains Mellum2.1 is an open 12B MoE coding agent model with 2.5B active parameters and 47.0% SWE-bench Verified.
Google Cloud’s Gemini agent is one enterprise agent routing work across Gemini and Claude models, with spend caps.
Claude Haiku 5.5 brings 1M context, adjustable effort, and $0.10 input pricing to high-volume subagent and browser workloads.
Unsloth details how Studio scans model code, blocks flagged weights, inspects packages and sandboxes tools before anything ...
Liquid AI has released Open d1, two open-weight multimodal models in its d1 decision model family. d1-3B reads text and ...
NVIDIA's PivotOPD trains multi-turn agents to recover from pivotal mistakes and beats 13 baselines on ALFWorld, WebShop, ...
Reflection AI's Beam is a 501B open-weight MoE with 23B active parameters, 1M context, and Apache 2.0 weights.