Research

Notes, methods, and lab updates on capturing expert work as training material for frontier AI. Long-form when it warrants it, short when it doesn't.

Introducing VOICE-H.

We're releasing VOICE-H, a new human evaluation benchmark for TTS models: 300 raters, 100 quotes selected from 40,000 real interviews, nine models against the original human recordings, and the qualitative reasons behind every preference.