公開日:2026-06-09 00:00:07
動画タグ : AI,Artificial intelligence,machine learning,data annotation,data sourcing,data preparation,prelabeled datasets,natural language processing,ethical ai
動画概要 : Speech AI is everywhere – but can you trust the benchmarks used to evaluate it?
In the debut episode of The Data Layer by Appen, host Karla Heredia sits down with Eric Bezzam (Audio ML, Hugging Face) and Sergio Bruccoleri (VP of GenAI Operations, Appen) to unpack one of the most underappreciated problems in production AI: the gap between benchmark performance and real-world speech recognition.
Eric leads the Open ASR Leaderboard at Hugging Face – one of the most widely referenced rankings for speech models. In this episode, he explains why that leaderboard needed a significant update, and how a months-long collaboration with Appen’s data experts introduced a private evaluation track designed to combat “benchmaxxing”: the practice of over-optimizing models for known test sets rather than genuine generalization.
Sergio brings the practitioner’s perspective: what production-grade ASR actually looks like when models are deployed in noisy, accented, multilingual environments – and why the distance between a clean benchmark and a real deployment remains one of the hardest unsolved challenges in the field.
In this episode:
→ Why speech AI has become central again in the age of LLMs
→ Demo quality vs. production-grade ASR: vocabulary, latency, turn-taking, environment
→ What “benchmaxxing” is and why it threatens leaderboard integrity
→ How Appen’s private dataset, built over four months, shifted the rankings
→ Why Indian and Canadian English still show the largest gaps in top-ranked models
→ Whether English ASR is truly “solved” – and why the answer is more nuanced than most assume
→ What full-duplex voice systems, robotics, and accessibility mean for the next five years
Subscribe to The Data Layer by Appen for more conversations at the intersection of AI evaluation and real-world deployment.