Skip to content
VibekollenBETAVibekollen
BlogHugging Face

The Open ASR Leaderboard Adds Its First Global South Language

Article image or reusable cover for Hugging Face

Hugging Face has added two new language datasets to its Open ASR Leaderboard, a platform for comparing speech-to-text models.

The new datasets are in English and Hindi from India, designed to measure how well systems perform across different groups of people. Previous benchmarks showed only average values, but research shows that speech-to-text systems often perform significantly worse for certain demographics—for example, approximately twice as poorly for Black speakers than white speakers. The new Monsoon dataset is designed to reveal such disparities by varying across geography, age, gender, device type, and many other factors. Instead of recording from a small number of speakers in long sessions, data was collected from nearly 5,000 different speakers from hundreds of districts across India, many with only a single recording each.

A model that scores well on the Open ASR Leaderboard gets adopted and iterated on, while capabilities the leaderboard does not measure tend not to improve.
Verbatim from the article at Hugging Face
Read the full story at Hugging Face →

Vibekollen prepared this summary with AI from the original publication. The content belongs to Hugging Face.

More to read