Cartesia's Sonic-3.6 has claimed the top spot on both Artificial Analysis speech leaderboards for streaming text-to-speech, marking a shift away from the dominant Transformer architectures. The significance lies in independent verifiability: on one leaderboard, models are cloned onto the same eight reference voices to isolate synthesis engine efficiency from voice catalogs. In this setup, Sonic-3.6 outperforms its predecessor (Sonic-3.5) and ElevenLabs' Eleven v3.
State Space Architecture and Sub-90ms Latency
The technical core of Sonic-3.6 is the adoption of state space models instead of Transformers, enabling a time-to-first-audio below 90 milliseconds. This is critical for real-time voice agents, where conversational fluidity depends on rapid audio generation. Available in beta and via API, the release natively supports 44 languages and handles confirmation codes and heteronyms correctly without preprocessing.

Cartesia Sonic-3 Tested: 90ms Voice AI, 10-Sec Clone | ThePlanetTools.ai — https://theplanettools.ai/tools/cartesia
Artificial Analysis Leadership and Market Context
The top ranking on Artificial Analysis solidifies Cartesia's strategy, which already occupied both ends of the voice pipeline with Sonic-3.5 and Ink-2. The model enters a market where latency and naturalness are the primary differentiators, surpassing the limitations of Transformer-based solutions that often struggle to balance quality and response speed in streaming.

AI-generated comment
AI-generated comment
AI-generated comment
AI-generated comment