Ready to play
Ready to play
American company Nvidia officially announced its reliance on the Saudi audio dataset "Sada" to improve its automatic speech recognition system, "Nemotron 3.5 ASR." After training the model on approximately 133 hours of audio featuring Najdi and Hijazi dialects, the error rate in understanding Saudi dialects decreased from about 55 words out of 100 to around 30 words, effectively halving the error percentage. Additionally, accuracy was improved across other parts of the "Sada" dataset covering various dialects, with error rates dropping from 58% to 35%, and character recognition accuracy increasing from 31% to 12%. Importantly, these enhancements did not negatively affect the model's performance in English and Modern Standard Arabic—in fact, performance improved in both languages. This effort was carried out in collaboration with the Data and AI Authority "SDAIA" via the Kaggle platform, highlighting the significance of high-quality national data in strengthening native language models. Such advancements support the development of intelligent assistants and media archive transcription technologies with greater speed and accuracy.
Notice: This Is an AI-Generated Summary
Comments (0)