Google's Gemini 3.8 Flash TTS and Flash-Lite TTS add prompt-based voice design, 2,000+ voices, and 100+ language support.
Contrastive-LM's CLM-8B scores agent actions instead of generating text, running up to 9× faster than TypeSafe's Jev zero-shot.
NVIDIA's Nemotron 3 Diarization is a 100M-parameter open-weight model that tracks 8 overlapping speakers in real-time streaming audio.
OpenAI launches GPT-6 Sol and Luna, lower-cost models priced from $0.10 per million input tokens with caching upgrades.
Voice input on phones has been solved for years. What has not been solved is the output. Speak into most dictation tools and ...
Anthropic has released Claude Opus 5.5, the first model in its new Claude 5.5 family. The team states it performs at the ...
Nokia's open-source AnyJev turns open LLMs into calibrated decision models with no training, lifting automatable traffic 6.8x ...
NVIDIA AI Releases SoL-Pi that uses auto-research loops to cut agent token traffic up to 49% and API cost about 33%.
Compare 7 voice cloning APIs, including ElevenLabs, Cartesia, and Inworld, on similarity, consent checks, licensing, and 2026 ...
SpaceXAI releases Grok Voice Transcribe 2.0, a speech-to-text API claiming 2x accuracy over 1.0 at $0.10 per hour.
TypeSafe AI's Jev returns typed decisions with calibrated probabilities instead of text, at $0.042 per 1M input tokens.
AWS releases Strands harness, an open-source Apache 2.0 agent harness reporting 28% lower token cost at comparable accuracy.