Research
Notes, methods, and lab updates on capturing expert work as training material for frontier AI. Long-form when it warrants it, short when it doesn't.
We're releasing VOICE-H, a new human evaluation benchmark for TTS models: 300 raters, 100 quotes selected from 40,000 real interviews, nine models against the original human recordings, and the qualitative reasons behind every preference.
No posts in this filter yet.