Contrastive-LM's CLM-8B scores agent actions instead of generating text, running up to 9× faster than TypeSafe's Jev zero-shot.
Google's Gemini 3.8 Flash TTS and Flash-Lite TTS add prompt-based voice design, 2,000+ voices, and 100+ language support.
NVIDIA's Nemotron 3 Diarization is a 100M-parameter open-weight model that tracks 8 overlapping speakers in real-time streaming audio.
OpenAI launches GPT-6 Sol and Luna, lower-cost models priced from $0.10 per million input tokens with caching upgrades.
Nokia's open-source AnyJev turns open LLMs into calibrated decision models with no training, lifting automatable traffic 6.8x ...
Voice input on phones has been solved for years. What has not been solved is the output. Speak into most dictation tools and ...
Anthropic has released Claude Opus 5.5, the first model in its new Claude 5.5 family. The team states it performs at the ...
NVIDIA AI Releases SoL-Pi that uses auto-research loops to cut agent token traffic up to 49% and API cost about 33%.
Linkup Research releases SPARSEUP, an open-source 149M-parameter sparse embedding model scoring 56.4 nDCG@10 on BEIR-13 under ...
Compare 7 voice cloning APIs, including ElevenLabs, Cartesia, and Inworld, on similarity, consent checks, licensing, and 2026 ...
Compare 11 open-source agent harnesses for local LLMs, including OpenCode, Pi, Goose, Cline, OpenHands, Aider, and Codex CLI.
Kyutai's Voice of Reason uses reinforcement learning to lift GLM-4-Voice from 27.3% to 77.1% on spoken GSM8K math.