We have added five TTS models to VOICE-H, taking it to fifteen, and re-ranked them all on a new round of comparisons. Cartesia's Sonic 3.6 takes the top place from Gemini 3.1 Flash, though only by a thin margin. Speechify's Simba 3.2 arrives third, at a sixth of the price of comparable models.
What separates the field
Every comparison in VOICE-H ends with a rater explaining their choice in their own words, and most of those explanations contain either a compliment, a complaint, or both. These produce a ratio that matches the leaderboard almost exactly.
| Model | Praised | Criticised | Ratio |
|---|---|---|---|
| Sonic 3.6 | 18.4 | 12.1 | 1.52 |
| Gemini 3.1 Flash | 16.5 | 20.0 | 0.82 |
| Simba 3.2 | 11.9 | 15.0 | 0.80 |
| Eleven v3 | 11.4 | 19.5 | 0.58 |
| Aura 2 | 9.9 | 21.1 | 0.47 |
| Flux | 8.9 | 29.3 | 0.30 |
| Eleven v2 | 6.0 | 35.7 | 0.17 |
| Speechmatics | 4.7 | 41.8 | 0.11 |
Mentions per hundred appearances, English.
Sonic 3.6 is the only model in the study that raters praise more often than they criticise. Speechmatics, at the bottom of the table, receives nearly nine complaints for every compliment. The order this produces is close enough to the order the Elo produces that the two are hard to tell apart.
Gemini 3.1 Flash is praised almost as often as Sonic 3.6, 16.5 against 18.4. Where they differ more strongly is how often they are criticised, 20.0 against 12.1. This gap is what leaves Gemini in second place.
Sonic 3.6 vs Sonic 3.5
The most common complaint about Sonic 3.5 in the first round was that it sounded wrong for the material. Raters called it seductive, sultry and breathy, on completely mundane passages. That complaint has largely gone away with Sonic 3.6. Breathy characteristics appeared in 3 of every hundred comparisons involving 3.5, and now only 0.9 of every hundred involving 3.6. A significant improvement.
The prompts we use are verbatim speech, with frequent half-finished words, self-corrections and stray fillers, and most models are marked down for how they read them. Sonic 3.6 is praised for it more than three times as often as it is criticised, which is something that no other model manages.
“These both sounded very natural, handled punctuation very well. Both models were capable of speeding up and slowing down, which gives a more human sound to the text. But I felt B [Sonic 3.6] sounded excellent and was more natural and humanlike than A [Voxtral Mini], especially around punctuation pauses, and lowering or raising and speeding up the voice at the point of something that is meant to be read as an aside.”
Voxtral Mini vs Sonic 3.6
Only sixty-six Elo separates Sonic 3.6 from Sonic 3.5, but at the top of the leaderboard that's difference between fourth place and first.
Nothing to complain about
Raters mention Simba 3.2, favourably or otherwise, significantly less than other models. They just don't seem to have much to say about it. It is almost never faulted.
Read its comments and the vocabulary is mild. Pleasant. Clear. Believable. Its one recurring complaint is speed.
“Lacking excitement. It seems to prioritise speed instead of highlighting the emotions. Pronunciation is great.”
Sonic 3.6 vs Simba 3.2
Nothing about Simba stands out, but against a model with a particular flaw that is enough to win. It finishes third, between two models that cost about six times as much.
Price appears to say little about how a model ranks. The next cheapest after Simba finishes last. The two dearest are priced identically and finish six places apart.
Newer is not necessarily better
Flux is Deepgram's newer model and it places four below their Aura 2 model. On every complaint Aura 2 itself attracted, Flux attracted more.
| Complaint | Aura 2 | Flux |
|---|---|---|
| Sterile, robotic, stock | 8.8 | 10.1 |
| Flat, monotone, no emotion | 5.0 | 8.0 |
| Like reading a script | 1.0 | 1.8 |
| Rushed, no pauses | 4.2 | 5.3 |
Mentions per hundred appearances, English.
Aura 2's weakness in English is that it sounds flat and generic.
“Strongly prefer A [Human Clip]. The speaker's accent gave some authenticity. B [Aura 2] sounded stock and sterile.”
Human Clip vs Aura 2
Flux draws the same criticism more often, and rushing on top of it. It reads at an ordinary speed and simply does not stop.
“Option B [Flux] was very rushed and did not acknowledge grammar, no full stops were adhered to, making it a very quick message, lowering the emotion behind what they were saying.”
Eleven v2 vs Flux
Aura 2 needs more expressive range. Flux needs the same, and to incorporate natural pauses.
The ranking differs by language
Round two added two models to French and Spanish. Along with taking first place in English, Sonic 3.6 also took first in Spanish and second in French. Inworld TTS-2 came third in Spanish and eighth in French.
Spanish raters call Inworld robotic in 4 of every hundred comparisons, among the lowest rates in that language. French raters do so in 14, among the highest. Its intonation is what Spanish comments praise and what French comments fault. In French it still scores above average for accuracy, but below it for naturalness.
| Model | English | French | Spanish |
|---|---|---|---|
| Sonic 3.6 | 1 | 2 | 1 |
| Gemini 3.1 Flash | 2 | 3 | 2 |
| Sonic 3.5 | 4 | 4 | 4 |
| Human recording | 5 | 7 | 6 |
| Inworld TTS-2 | 6 | 8 | 3 |
| Eleven v3 | 7 | 1 | 8 |
| Aura 2 | 8 | 10 | 7 |
| Voxtral Mini | 9 | 9 | 11 |
| GPT-4o Mini TTS | 10 | 11 | 10 |
| Pocket TTS | 11 | 6 | 5 |
| Eleven v2 | 13 | 5 | 9 |
| Lightning V3.1 | 14 | 12 | 12 |
The twelve models that were rated in all three languages.
The pattern is not new. Eleven v3 finished first in French and fifth in English in the first round, and it still leads in French. Pocket TTS, an open model of a hundred million parameters that runs on a laptop, is eleventh in English and fifth in Spanish.
The three at the top hold across all three languages. Below them the order is language-specific, and four models move by five places or more.
How this round was run
The first round measured every pairing among ten models. The second added five more and fielded only the comparisons that had not been made yet. The two rounds are then fitted together as one set of votes.
That only works if a vote in the first round means the same as a vote in the second. Each language re-measured a sampling of first-round pairings months later with different people, and the results held, which rules out any shift larger than a few points.
Where this leaves things
Twenty Elo separates the top two models in English, which is too close to call confidently. Below that, neither price nor recency predicts placement: the cheapest model with a price tag finishes third, and Deepgram's newer release finishes behind the one it replaces. Nor does a ranking in one language hold in another. A model that comes third in Spanish comes eighth in French. The first round produced a leader that stood clear of the rest. This one is less clear-cut.
Explore the full interactive results: head-to-head records, preference distributions, and more →