Robust ASR
- 1 Wifaq 43.23
- 2 ResonateNile University 50.29
- 3 ThakaThaka, Advanced AI & Information Technology 50.41
A shared task on five speech-processing challenges for the Arab world’s living dialects , including recognition, identification, synthesis, translation, and spoken understanding.
Arabic speech systems still struggle: Many current systems perform reasonably well on Modern Standard Arabic (MSA) and clean speech from dominant dialects, but degrade under realistic conditions such as background noise, low-bandwidth audio, mixed dialects, and sub-country regional variation.
Building on NADI 2025’s spoken dialect work, NADI 2026 broadens the focus with three main speech-focused families of tasks: robust dialectal ASR for real-world conditions, spoken dialect identification under cross-domain conditions, and dialectal Arabic text-to-speech as a new generative component — alongside speech translation and spoken language understanding.
Together, these tasks benchmark discriminative and generative Arabic speech systems to advance robust, inclusive technologies reflecting the Arab world’s linguistic diversity.
Each task ships with new blind test data. Baselines, evaluation scripts, and submission links released with the data on June 16, 2026.
This table summarizes the main resources for each task, including notebooks, training data, development data, and the relevant platform or submission path. Task 3 TTS follows a different evaluation and submission flow, so please check its submission instructions button.
| Task | Baselines | Training data | Development data | Test data | Evaluation | ||||
|---|---|---|---|---|---|---|---|---|---|
| Notebook | #hours | #utterances | Dataset | #hours | #utterances | Dataset | Dataset | Platform | |
Please contact the shared task chairs for any issues with affiliation information, or other issues.
Combined NADI shared-task milestones and ArabicNLP 2026 conference deadlines. Source: arabicnlp2026.sigarab.org.
| Date | Milestone | Status |
|---|
Fill out the registration form to receive access to training and development data, baseline systems, and submission links.
Registration form →Submit via CodaBench for ASR & SID where suitable, and Hugging Face Spaces for large TTS audio submissions. SID hidden test runs through a private platform.
View baselines →Submit a system description paper by August 22. Document external data, pretrained models, preprocessing, and decoding settings clearly.
Paper guidelines →For questions contact the organizers at nadisharedtask@gmail.com.
The system description paper should let another researcher:
The paper will be included in the The Fourth Arabic Natural Language Processing Conference (ArabicNLP 2026) proceedings. Please familiarize yourself with the general shared task paper requirements for the conference.
The paper is expected to be up to 4 pages of content, plus unlimited references and appendices; final versions of the paper will be given one additional page of content (up to 5 pages) so that reviewers’ comments can be taken into account. Please, note that the review process is not double-blind, so anonymity is not required.
Paper submissions must use the official ACL style templates, which are available here (Latex). Please follow the paper formatting guidelines, general to "*ACL" conferences available here.
note: use the acl_latex.tex template but do not use acl_lua_latex.tex due to poorer Arabic language support in LuaLaTex.
Authors may not modify these style files or use templates designed for other conferences.
A common structure for system description papers is:
Title should be as follows: your_team_name at NADI 2026 shared task: your_own_title. For example, UBC at NADI 2026 shared task: Multitask learning for Arabic Dialect Identification
four/five sentences highlighting your approach and key results.
¾ a page expanding on the abstract mentioning key background such as why the task is challenging for current modeling techniques and why your approach is interesting/novel.
review of the data you used to train your system. Be sure to mention the size of the training, validation and test sets that you’ve used, and the label distributions, as well as any tools you used for preprocessing data.
a detailed description of how the systems were built and trained. If you’re using a neural network, were there pre-trained embeddings, how was the model trained, what hyperparameters were chosen and experimented with? How long did the model take to train, and on what infrastructure? Linking to source code is valuable here as well, but the description should be able to stand alone as a full description of how to reimplement the system. While other paper styles include background as a separate section, it’s fine to simply include citations to similar systems which inspired your work as you describe your system.
a description of the key results of the paper. If you have done extra error analysis into what types of errors the system makes, this is extremely valuable for the reader. Unofficial results from after the submission deadline can be very useful as well.
general discussion of the task and your system. Description of characteristic errors and their frequency over a sample of development data. What would you do if you had another 3 months to work on it?
a restatement of the introduction, highlighting what was learned about the task and how to model it.
@inproceedings{Sullivan-etal-2026-nadi,
title = "{NADI-2026: The Second Multidialectal {A}rabic Speech Processing Shared Task}",
author = "Sullivan, Peter and
Talafha, Bashar and
Ashraf, Ahmed and
Bougares, Fethi and
Elleuch, Haroun and
Zhang, Chiyu and
Elmadany, AbdelRahim and
Mohamed, Youssef and
Mdhaffar, Salima and
Est{\`e}ve, Yannick and
Elhoseiny, Mohamed and
Luqman, Hamzah and
Habash, Nizar and
Abdul-Mageed, Muhammad",
booktitle = "Proceedings of the Fourth Arabic Natural Language Processing Conference (ArabicNLP 2026)",
year = "2026",
address = "Budapest, Hungary",
publisher = "Association for Computational Linguistics",
}
@inproceedings{abdul-mageed-etal-2020-nadi,
title = "{NADI} 2020: The First Nuanced {A}rabic Dialect Identification Shared Task",
author = "Abdul-Mageed, Muhammad and
Zhang, Chiyu and
Bouamor, Houda and
Habash, Nizar",
booktitle = "Proceedings of the Fifth Arabic Natural Language Processing Workshop",
month = dec,
year = "2020",
address = "Barcelona, Spain (Online)",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2020.wanlp-1.9",
pages = "97--110",
}
@inproceedings{abdul-mageed-etal-2021-nadi,
title = "{NADI} 2021: The Second Nuanced {A}rabic Dialect Identification Shared Task",
author = "Abdul-Mageed, Muhammad and
Zhang, Chiyu and
Elmadany, AbdelRahim and
Bouamor, Houda and
Habash, Nizar",
booktitle = "Proceedings of the Sixth Arabic Natural Language Processing Workshop",
month = apr,
year = "2021",
address = "Kyiv, Ukraine (Virtual)",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2021.wanlp-1.28",
pages = "244--259",
}
@inproceedings{abdul-mageed-etal-2022-nadi,
title = "{NADI} 2022: The Third Nuanced {A}rabic Dialect Identification Shared Task",
author = "Abdul-Mageed, Muhammad and
Zhang, Chiyu and
Elmadany, AbdelRahim and
Bouamor, Houda and
Habash, Nizar",
booktitle = "Proceedings of the The Seventh Arabic Natural Language Processing Workshop (WANLP)",
month = dec,
year = "2022",
address = "Abu Dhabi, United Arab Emirates (Hybrid)",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.wanlp-1.9",
pages = "85--97",
}
@inproceedings{abdul-mageed-etal-2023-nadi,
title = "{NADI} 2023: The Fourth Nuanced {A}rabic Dialect Identification Shared Task",
author = "Abdul-Mageed, Muhammad and
Elmadany, AbdelRahim and
Zhang, Chiyu and
Nagoudi, El Moatez Billah and
Bouamor, Houda and
Habash, Nizar",
editor = "Sawaf, Hassan and
El-Beltagy, Samhaa and
Zaghouani, Wajdi and
Magdy, Walid and
Abdelali, Ahmed and
Tomeh, Nadi and
Abu Farha, Ibrahim and
Habash, Nizar and
Khalifa, Salam and
Keleg, Amr and
Haddad, Hatem and
Zitouni, Imed and
Mrini, Khalil and
Almatham, Rawan",
booktitle = "Proceedings of ArabicNLP 2023",
month = dec,
year = "2023",
address = "Singapore (Hybrid)",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.arabicnlp-1.62",
doi = "10.18653/v1/2023.arabicnlp-1.62",
pages = "600--613",
}
@inproceedings{abdul-mageed-etal-2024-nadi,
title = "{NADI} 2024: The Fifth Nuanced {A}rabic Dialect Identification Shared Task",
author = "Abdul-Mageed, Muhammad and
Keleg, Amr and
Elmadany, AbdelRahim and
Zhang, Chiyu and
Hamed, Injy and
Magdy, Walid and
Bouamor, Houda and
Habash, Nizar",
editor = "Habash, Nizar and
Bouamor, Houda and
Eskander, Ramy and
Tomeh, Nadi and
Abu Farha, Ibrahim and
Abdelali, Ahmed and
Touileb, Samia and
Hamed, Injy and
Onaizan, Yaser and
Alhafni, Bashar and
Antoun, Wissam and
Khalifa, Salam and
Haddad, Hatem and
Zitouni, Imed and
AlKhamissi, Badr and
Almatham, Rawan and
Mrini, Khalil",
booktitle = "Proceedings of the Second Arabic Natural Language Processing Conference",
month = aug,
year = "2024",
address = "Bangkok, Thailand",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.arabicnlp-1.79/",
doi = "10.18653/v1/2024.arabicnlp-1.79",
pages = "709--728",
}
@inproceedings{talafha-etal-2025-nadi,
title = "{NADI} 2025: The First Multidialectal {A}rabic Speech Processing Shared Task",
author = "Talafha, Bashar and
Toyin, Hawau Olamide and
Sullivan, Peter and
Elmadany, AbdelRahim A. and
Juma, Abdurrahman and
Djanibekov, Amirbek and
Zhang, Chiyu and
Alshehhi, Hamad and
Aldarmaki, Hanan and
Jarrar, Mustafa and
Habash, Nizar and
Abdul-Mageed, Muhammad",
editor = "Darwish, Kareem and
Ali, Ahmed and
Abu Farha, Ibrahim and
Touileb, Samia and
Zitouni, Imed and
Abdelali, Ahmed and
Al-Ghamdi, Sharefah and
Alkhereyf, Sakhar and
Zaghouani, Wajdi and
Khalifa, Salam and
AlKhamissi, Badr and
Almatham, Rawan and
Hamed, Injy and
Alyafeai, Zaid and
Alowisheq, Areeb and
Inoue, Go and
Mrini, Khalil and
Alshammari, Waad",
booktitle = "Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks",
month = nov,
year = "2025",
address = "Suzhou, China",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.arabicnlp-sharedtasks.99/",
doi = "10.18653/v1/2025.arabicnlp-sharedtasks.99",
pages = "720--733",
ISBN = "979-8-89176-356-2",
}
Registration opens May 16, 2026. Training and development data, baseline systems, and evaluation scripts land June 16. Blind test data ships July 20.