NI Global logo
NI GLOBAL

Noble • Iconic • Unstoppable

Back to Insights
AISpeech RecognitionMachine LearningResearchTechnology

Indic DiarBench: Advancing Joint ASR and Speaker Diarization for Indian Languages

adminAugust 13, 20262 min read
Indic DiarBench: Advancing Joint ASR and Speaker Diarization for Indian Languages

Indic DiarBench: Advancing Joint ASR and Speaker Diarization for Indian Languages




While speech recognition models excel for single-speaker dictation, real-world conversations pose a challenge. Rapid turn-taking, overlapping voices, and code-mixing disrupt accuracy, especially in complex linguistic landscapes like India.




To address this gap, Sarvam AI and AI4Bharat have introduced Indic DiarBench. This is the first open benchmark dataset specifically designed for joint Automatic Speech Recognition (ASR) and speaker diarization across all 22 scheduled Indian languages.




Historically, models have processed diarization (identifying who is speaking) and ASR (transcribing what is being said) independently. This separation often results in accurate transcriptions being attributed to the incorrect speaker, undermining the overall utility of the system. Indic DiarBench resolves this by facilitating the simultaneous evaluation of both components on natural, unscripted speech, ensuring that a model's performance reflects true conversational dynamics.




Key Features of the Benchmark


▸Comprehensive Scale: Contains approximately 108 hours of natural, multi-speaker audio featuring 2 to 9 participants per recording.


▸Linguistic Diversity: Captures 485 unique speakers from 189 districts, representing a wide range of urban and rural dialects and educational backgrounds.


▸Realistic Acoustic Environments: Comprises near-field meetings (~53 hours), far-field meetings with complex reverberation profiles and background noise (~27 hours), and unpredictable "in-the-wild" media sourced from public broadcasts and YouTube (~28 hours).


▸High Speech Overlap: Deliberately includes rapid backchanneling, interruptions, and concurrent speaking (capturing up to four simultaneous voices) to accurately mirror genuine human interaction.


▸Code-Mixing Support: Transcripts are meticulously annotated in both native Indic scripts and Romanized English to reflect authentic bilingual speech patterns, ensuring models can handle the fluid transition between languages.


▸Rigorous Human Verification: Every word, timestamp, and speaker label underwent a multi-stage human review process to guarantee high-quality annotations, far surpassing auto-generated datasets.




By enabling evaluation through rigorous metrics such as DER (Diarization Error Rate), cpWER (concatenated minimum-permutation Word Error Rate), and WDER (Word Diarization Error Rate), Indic DiarBench provides developers with the robust framework needed to build highly reliable conversational AI systems.




This initiative represents a critical step forward in localizing AI infrastructure. The dataset is fully open-source and available for researchers and developers to explore on Hugging Face.




#IndicDiarBench #SarvamAI #AI4Bharat #SpeechRecognition #ConversationalAI #ASR #IndicLanguages #OpenSourceAI #MachineLearning #TechIndia #HuggingFace #DataScience

#Indic DiarBench#ASR#Speaker Diarization#Sarvam AI#AI4Bharat#Indian Languages#Conversational AI#Speech Recognition#Open Source AI#Machine Learning#Hugging Face#AI Research