Human evals · July 2026

VOICE-H

A human evaluation benchmark for text-to-speech models. 300 raters from the Askable Labs panel judged 4,500 blind pairwise comparisons across 100 prompts drawn from an Askable Labs dataset of over 40,000 real user interviews: nine TTS models against the original human recordings, rated for naturalness, accuracy, and emotion/tone, with the reasoning behind every preference. Read the full write-up and findings in the research post.

Leaderboard

Elo from an ordinal Bradley-Terry fit weighted by preference strength. Whiskers are 95% bootstrap confidence intervals. The human recordings are the red dot.

Head-to-head

Each cell is the row system’s win rate against the column system over decisive (non-tie) votes. Amber rows win; indigo rows lose.

How strongly raters preferred

Distribution of the five preference labels. Sides were randomised per comparison. A small number of the 4,500 comparisons (26) were removed by rater-quality filtering, so the bars sum to 4,474.

Price vs performance

Overall Elo against list-price cost per prompt (the eval’s average prompt is 645 characters, ≈49 seconds of audio). Pocket TTS is open-source and self-hostable, so its marginal cost is set to zero. The human recordings have no API price and appear as a reference line.

Who rated

The rater pool, from the Askable Labs panel. No rater saw more than 15 comparisons, all for a single prompt.

Methodology

100 prompts were curated from over 40,000 real Askable user interviews: 10 seconds to three minutes of contiguous, emotional speech, manually reviewed for transcription accuracy. Nine TTS models plus the original human recording give 45 unique pairs per prompt; each rater judged 15 randomised comparisons for one prompt, rating each sample for naturalness, accuracy, and emotion/tone, then giving an overall preference on a five-point scale with free-text reasoning. Elo scores use ordinal Bradley-Terry preference with weighting, taking into account the strength of the preference, with bootstrap confidence intervals. The full methodology, findings, and rater quotes are in the research post.