Accepted African Datasets
Accepted Datasets
Chichewa Text-to-SQL Benchmark Dataset
Presenter Name: John Emeka Eze
Dataset Domain: Agriculture, Education, Language & Culture | Langue et culture, Infrastructure
Video Presentation: https://drive.google.com/open?id=18Mp1HnlZg8Rp4yuzAIle0XTDRcSAr-gb
Dataset Description: The Chichewa Text-to-SQL dataset is a curated benchmark designed to support research in natural language interfaces for low-resource African languages. It consists of 400 natural language–SQL query pairs, where each query is written in Chichewa and mapped to structured SQL over a unified relational database. The dataset covers practical domains such as agriculture, commodity prices, population statistics, and food security, reflecting real-world information needs in Malawi and similar contexts. It was developed through careful translation and validation to ensure semantic consistency between queries and database schemas. This dataset enables the evaluation and development of machine learning models, particularly large language models, for structured query generation in data-constrained and linguistically underrepresented environments.
Nigeria
Cactus_dataset_V2
Presenter Name: Adel BENALI
Dataset Domain: Agriculture
Citation: BENALI, A., Fourati, R., & JDEY, I. (2026). Cactus Disease dataset V2 (Version 2) [Data set]. Zenodo.
Dataset Description: This repository contains two image datasets of cacti designed for scientific research in plant disease classification and segmentation using machine learning. All images are labeled and organized in folders according to their respective classes.
Tunisia
OHADA-CCJA Court Decisions Corpus: Case Law of the Common Court of Justice and Arbitration for African Legal NLP
Presenter Name: Foutse Yuehgoh
Dataset Domain: Legal & Economic Development / Law & Governance
Video Presentation: https://drive.google.com/open?id=1eSSDrmBpEK7UPnN7xxD1WNqvVvQ8a6wz
Dataset Description: A curated corpus of 4,059 court decisions from the Cour Commune de Justice et d'Arbitrage (CCJA), the supranational court of the Organisation pour l'Harmonisation en Afrique du Droit des Affaires (OHADA). OHADA harmonizes business law across 17 African member states: Benin, Burkina Faso, Cameroon, Central African Republic, Chad, Comoros, Democratic Republic of Congo, Republic of Congo, Côte d'Ivoire, Equatorial Guinea, Gabon, Guinea, Guinea-Bissau, Mali, Niger, Senegal, and Togo. This dataset provides structured access to CCJA jurisprudence spanning over two decades (1997–2023), making it a unique resource for African legal NLP research.
France
UCCB: Uganda Cultural Context Benchmark
Presenter Name: Lameck Kavuma
Dataset Domain: Language & Culture | Langue et culture
Video Presentation: https://drive.google.com/drive/folders/1gIMKyaN27CgpGl9hecoiF3CsFTgQVd-X?usp=sharing
Full data card: https://huggingface.co/datasets/CraneAILabs/UCCB
GitHub repository: https://github.com/Crane-AI-Labs/UCCB
UK AI Safety Institute Inspect AI integration: https://github.com/UKGovernmentBEIS/inspect_evals/tree/main/src/inspect_evals/uccb
Dataset Description: The Uganda Cultural Context Benchmark (UCCB) is the first comprehensive question-answer dataset designed to evaluate the cultural understanding and reasoning abilities of Large Language Models (LLMs) concerning Uganda. It contains 1,039 expert-validated question-answer pairs across 24 cultural domains including folklore, traditional medicine, music, attire, slang, history, religion, and economy. The benchmark incorporates terminology from multiple indigenous Ugandan languages (Luganda, Runyankole, Acholi, Ateso, Lugbara, among others) embedded within English-language questions. UCCB was curated through a hybrid pipeline combining Wikipedia-sourced content with LLM-assisted generation, chunk-grounded quality scoring, and validation by 10 Ugandan annotators from across the country's four major regions. The dataset has been integrated into the UK AI Safety Institute's Inspect AI framework, marking the first African cultural benchmark in a sovereign government's AI evaluation infrastructure. Evaluation of five frontier LLMs revealed a 2.94-point spread on a 5-point scale, demonstrating significant variance in cultural knowledge across model families.
Uganda
African Urban Land Use and Land Cover Segmentation Dataset (AULC-13)
Presenter Name: Benayad Mohamed
Dataset Domain: Agriculture, Education, Environment | Environnement, Infrastructure
Dataset Description:
AfriUrban-13 is a high-resolution satellite image dataset designed for semantic segmentation of urban land use and land cover in African cities. The dataset contains manually annotated RGB satellite image tiles labeled into 13 classes: roads, vegetation, grass, trees, brown soil, bare land, buildings, solar panels, tracks, sidewalks, playgrounds, cars, and water. The dataset was developed to address the underrepresentation of African urban environments in existing global segmentation benchmarks. It captures the spatial complexity of semi-arid urban areas, including mixed land cover patterns, informal structures, distributed solar installations, and transportation infrastructure. AfriUrban-13 is intended to support the development and evaluation of deep learning models, including CNN-based architectures (e.g., U-Net) and transformer-based segmentation models (e.g., SegFormer). It can be used for applications in smart city planning, infrastructure development, renewable energy mapping, environmental monitoring, and sustainable urban development. By providing high-quality pixel-level annotations from an African context, this dataset aims to strengthen locally relevant AI research and contribute to more inclusive global geospatial benchmarks.
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
Morocco
AI Startups in Africa
Presenter Name: Chinasa T. Okolo
Dataset Domain: Industry
Video Presentation: https://drive.google.com/open?id=1AxSR4zUatstUbU3pTfUNyKyG3idQfMNp
Dataset Description: A list of AI startups in Africa with company name, sector/domain, location, customer type, year founded, team, stage, operating status, and brief description.
United States
Senegal Maternal Health QA dataset (SenMH-QA)
Presenter Name: Ertony Basilwango
Dataset Domain: Health | Santé, Language & Culture | Langue et culture
Video Presentation: https://drive.google.com/open?id=1ATVDQUqCR0eGbJDqHC9UbZ3s-0aV4eLg
Dataset Description: This dataset contains short audio recordings collected in community settings in Senegal. The recordings cover maternal health topics, including family planning, healthcare access, pregnancy practices, and cultural beliefs around maternal and reproductive health. Audio was captured using mobile devices or portable recorders in natural conversational conditions, and all transcriptions were manually verified by Wolof linguists, maternal health experts, and machine learning engineers. The dataset is intended for training and benchmarking domain-specific automatic speech recognition (ASR) models.
Senegal
AfricaBias-SW-FR: A Multilingual Bias Detection Dataset for Swahili and French
Presenter Name: Preston Osoro
Dataset Domain: Language & Culture | Langue et culture
Dataset Description: AfricaBias-SW-FR is a multilingual bias detection dataset containing 35,285 annotated sentences in Swahili and French, developed to address the near-total absence of bias benchmarks for African languages. Each sentence is labelled across four categories, neutral, stereotype, counter-stereotype, and derogation, spanning eight social domains, including household and care, livelihoods and work, governance, and health. The dataset draws from real Kenyan and Francophone African sources, including news, social media, encyclopedias, and government documents, and captures gender, religious, and cultural sensitivity attributes. It is the first publicly released large-scale gender and social bias detection dataset for Swahili, produced as a Gates Foundation-supported deliverable through the AfriLabs AI Accelerator programme.
Kenya
mghana_st
Ghana
Enugu Dialect Igbo language and cultural multimodal dataset
Presenter Name: Peter Emmanuel Olotuche
Dataset Domain: Education, Environment | Environnement, Language & Culture | Langue et culture
Dataset Description: This dataset contains multimodal data representing the Igbo language and culture from Enugu State, Nigeria. It includes text documents, images, audio recordings, and short videos capturing everyday conversations, cultural expressions, environmental contexts, and educational materials. The dataset is designed to support AI training, language preservation, and research in African languages and cultural knowledge.
Nigeria
African Plums Dataset
Presenter Name: FAMENI TAGNI ARMEL GABIN
Dataset Domain: Agriculture
Dataset Description: This dataset contains 4,507 annotated images of African plums collected from various fields in Cameroon. It is specifically designed for training and evaluating AI models in fruit quality assessment and defect detection. The images are categorized into six classes based on their defect type: bruised, cracked, rotten, spotted, unaffected, and unripe.
Cameroon
FITSIPIKA Malagasy Dataset (SOKAJY)
Presenter Name: Vatosoa Razafindrazaka
Dataset Domain: Language & Culture | Langue et culture
Dataset Description: A Part-of-speech corpus for Malagasy (28M Speakers) created to bridge the gap in low-resource NLP. It includes manual annotations for 4,693 sentences. Validated through peer-reviewed publication in Springer Nature, this dataset serves as a foundation for building ASR Malagasy and transcription systems.
License: CC BY-NC-SA 4.0