Google's Gemini 3.8 Flash TTS and Flash-Lite TTS add prompt-based voice design, 2,000+ voices, and 100+ language support.
NVIDIA's Nemotron 3 Diarization is a 100M-parameter open-weight model that tracks 8 overlapping speakers in real-time streaming audio.
Nokia's open-source AnyJev turns open LLMs into calibrated decision models with no training, lifting automatable traffic 6.8x ...
In the field of artificial intelligence, machine learning is a branch that uses data and algorithms to imitate human learning ...
Anthropic has released Claude Opus 5.5, the first model in its new Claude 5.5 family. The team states it performs at the ...
OpenAI launches GPT-6 Sol and Luna, lower-cost models priced from $0.10 per million input tokens with caching upgrades.
Voice input on phones has been solved for years. What has not been solved is the output. Speak into most dictation tools and ...
Kyutai's Voice of Reason uses reinforcement learning to lift GLM-4-Voice from 27.3% to 77.1% on spoken GSM8K math.
AWS releases Strands harness, an open-source Apache 2.0 agent harness reporting 28% lower token cost at comparable accuracy.
Alibaba's Qwen-Image-2.1 is a 7B open-weight model unifying image generation, editing, and native RGBA transparency under research license.
Compare 7 voice cloning APIs, including ElevenLabs, Cartesia, and Inworld, on similarity, consent checks, licensing, and 2026 ...
NVIDIA AI Releases SoL-Pi that uses auto-research loops to cut agent token traffic up to 49% and API cost about 33%.