Sarvam AI, in collaboration with AI4Bharat, has introduced Indic DiarBench, an open benchmark dataset to evaluate automatic speech recognition (ASR) and speaker diarisation across all 22 scheduled Indian languages.
The dataset contains 108 hours of natural, multi-speaker conversations featuring 485 unique speakers from 189 districts across urban and rural India, and is available on Hugging Face.
Sarvam AI, in collaboration with AI4Bharat, has introduced Indic DiarBench, an open benchmark dataset to evaluate automatic speech recognition (ASR) and speaker diarisation across all 22 scheduled Indian languages. The dataset contains 108 hours of natural, multi-speaker conversations featuring 485 unique speakers from 189 districts across urban and rural India, and is available on Hugging Face.