SIMBig 2026 will host a Shared Task focused on advancing speech technologies and resources for Indigenous languages, with an initial emphasis on Puno Quechua. The competition aims to encourage Natural Language Processing (NLP) to develop reproducible systems and open resources that contribute to the preservation and technological inclusion of Indigenous languages.
This shared task seeks to:
- Advance speech recognition technologies for Indigenous languages.
- Promote reproducible and open research.
- Encourage collaboration between NLP and Indigenous language communities.
- Increase the availability of open resources for indigenous languages.
- Strengthen the language research ecosystem in South America.
- The organisers welcome participation from researchers, students, industry practitioners, and Indigenous communities interested in developing technologies that support linguistic diversity and language preservation.
The shared task will use the Puno Quechua Speech dataset as the official benchmark for the baseline task. Participants will be encouraged to submit reproducible methods, publish their code, and compare results using common evaluation metrics.
Tasks
Task 1: Puno Quechua ASR Baseline
- Build an Automatic Speech Recognition (ASR) system using the Common Voice Puno Quechua datasets (Scripted Speech, Spontaneous Speech).
- Baseline model and evaluation scripts are provided (link).
- Objective: Improve transcription accuracy (i.e., WER and CER) while ensuring reproducibility.
Task 2: ASR for Any Indigenous Language of South America
- Participants may develop ASR systems for any Indigenous or native language that does not yet have an ASR system.
- Teams should provide the datasets and splits they used (train/dev/test), where the test dataset must be at least 30 minutes.
- Objective: Foster multilingual speech technologies beyond Puno Quechua and encourage broader participation from the research community.
Task 3: Educational/Literacy Resource based on Language Datasets
- Develop a resource that helps students, researchers, or the general public learn, listen, or explore indigenous language(s), or guides users in reusing those datasets.
- Possible submissions include interactive websites, tutorials, notebooks, visualisations, educational videos, and/or learning platforms.
- Objective: Improve accessibility and promote the use of Indigenous language resources in education and research.
Dataset
The official dataset for Task 1 will be the Common Voice Puno Quechua datasets (Scripted Speech, Spontaneous Speech), providing a common benchmark for system comparison. The baseline implementations and evaluation protocols are released on GitHub (link).
Awards*
- Task 1: $100 USD
- Task 2: $100 USD
- Task 3: $50 USD
In addition, the winning teams of each task will receive free registration to the SIMBig 2026.
* to be confirmed.
Important Dates
The shared task will be organised as part of the SIMBig 2026, to be held in Arequipa, Peru. All times (AoE):
- 10/08/2026 - Registration opens (form)
- 24/08/2026 - Release of datasets (Scripted Speech, Spontaneous Speech)
- 16/09/2026 - Open orientation hour
- 21/09/2026 - Registration closes
- 30/09/2026 - Release of test dataset for Task 1
- 18/10/2026 - Submission deadline for final results and system description paper
- 28-30/10/2026 - Results announcement during SIMBig 2026
In addition, the winning teams of each task will receive free registration to the SIMBig 2026.
Submission
Task1 and Task 2: System’s Predicted Transcriptions
- Once we release the test dataset (audio only), teams will have 1 week to submit their system’s predicted transcriptions. Submission should take the form of a zip file containing a TSV file (1 per language being attempted), e.g., qxp.tsv, and it should have two columns: the first is the name of the audio file, and the second is the predicted transcription.
- For Task 2, each team should provide their own splits (train/dev/test), where test must be of at least 30 minutes.
- The team with the best performance on each task will be asked to submit their model and an inference script so we can reproduce the results.
All Tasks: System Description Paper
Each submission will be evaluated by at least two independent reviewers based on originality, technical quality, significance, and relevance to the shared task. Additionally, all manuscripts will be screened using plagiarism detection software to ensure academic integrity and originality. Finally, the winners of each task, after addressing reviewers' feedback, will be published in the Springer Communications in Computer and Information Science - CCIS Book Series.
Organisers
Programme Chair:
- Elwin Huaman, University of Cambridge, UK.
- Johanna Cordova, ERTIM at Inalco, France.
- Juan Antonio Lossio-Ventura, National Institutes of Health, USA.
- Hugo Alatrista-Salas, De Vinci Research Centre, France.
Programme Committee:
- Johanna Cordova, ERTIM at Inalco, France.
- Elwin Huaman, University of Cambridge, UK.