BottleCap AI's ThinkingCap-Qwen3.8-27B cuts thinking tokens 37.2% across 12 benchmarks with only 0.86pp accuracy loss. Drop-in for vLLM.
Contrastive-LM's CLM-8B scores agent actions instead of generating text, running up to 9× faster than TypeSafe's Jev zero-shot.
OpenAI launches GPT-6 Sol and Luna, lower-cost models priced from $0.10 per million input tokens with caching upgrades.
Google's Gemini 3.8 Flash TTS and Flash-Lite TTS add prompt-based voice design, 2,000+ voices, and 100+ language support.
Anthropic has released Claude Opus 5.5, the first model in its new Claude 5.5 family. The team states it performs at the ...
Voice input on phones has been solved for years. What has not been solved is the output. Speak into most dictation tools and ...
NVIDIA's Nemotron 3 Diarization is a 100M-parameter open-weight model that tracks 8 overlapping speakers in real-time streaming audio.
NVIDIA AI Releases SoL-Pi that uses auto-research loops to cut agent token traffic up to 49% and API cost about 33%.
Compare 7 voice cloning APIs, including ElevenLabs, Cartesia, and Inworld, on similarity, consent checks, licensing, and 2026 ...
SpaceXAI releases Grok Voice Transcribe 2.0, a speech-to-text API claiming 2x accuracy over 1.0 at $0.10 per hour.
Nokia's open-source AnyJev turns open LLMs into calibrated decision models with no training, lifting automatable traffic 6.8x ...
Compare 11 open-source agent harnesses for local LLMs, including OpenCode, Pi, Goose, Cline, OpenHands, Aider, and Codex CLI.