Keynotes, talks, and presentations on AI, language technology, and Tetun.
Labadain: Resources and Applications for Tetun Language Technology
Date: July 24, 2026
Event: NLP Group at the University of Melbourne
Venue: melbourne, Australia
This talk presents Labadain, a suite of resources and applications for Tetun language technology, including datasets, tools, and AI systems that enable inclusive digital access for Tetun speakers.
Labadain: Language Technology Resources and Tools for Tetun
Date: June 03, 2026
Event: Presenter at the 2nd ASEAN-NEXUS ECRs Enclave
Venue: Singapore
This presentation introduces resources and tools developed for the Tetun language, highlighting the latest achievements and ongoing initiatives of the Labadain project to the ASEAN NEXUS community.
This talk presents Labadain LIX-R361, an agentic AI system for Tetun that improves performance through LLM customization and supports use cases such as translation, news generation, and writing assistance, while highlighting opportunities and challenges for AI in Timor-Leste.
AI for Tetun: Building Timor-Leste's Inclusive Digital Future
Date: November 21, 2025
Event: Keynote Speaker at the TLNOG2 Conference
Venue: Dili, Timor-Leste
This talk explores how AI can support Tetun, a low-resource and official language of Timor-Leste, by enabling inclusive digital access through language technologies, datasets, and information retrieval systems.
Labadain: The Foundation of Tetun Language Technology
Date: November 20, 2025
Event: Invited Talk at the DEI–FECT–UNTL National Seminar
Venue: Dili, Timor-Leste
This talk introduces Labadain as the foundation of Tetun language technology, showcasing how datasets, tools, and AI systems enable inclusive digital access for Tetun speakers.
Conference Proceedings Talk on the Labadain Crawler Pipeline for LRLs
Date: May 22, 2024
Event: Conference proceedings talk at the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
Venue: Torino, Italy
This talk presents Labadain Crawler, a web-based data collection pipeline designed for low-resource languages, detailing its architecture, language processing components, and its application to building a high-quality Tetun text corpus.
Conference Proceedings Talk on Labadain-30k+ Dataset Construction
Date: May 20, 2024
Event: Conference proceedings talk at the 3rd Annual Meeting of the Special Interest Group on Under-resourced Languages at LREC-COLING 2024
Venue: Torino, Italy
This talk presents Labadain-30k+, a manually audited Tetun text dataset, outlining the data collection pipeline, quality control process, and key insights from content analysis to support NLP and information retrieval research in a low-resource language.