{"id":1214,"date":"2026-08-04T12:36:49","date_gmt":"2026-08-04T12:36:49","guid":{"rendered":"https:\/\/deeplearningindaba.com\/2026\/?page_id=1214"},"modified":"2026-08-04T13:02:22","modified_gmt":"2026-08-04T13:02:22","slug":"accepted-african-datasets","status":"publish","type":"page","link":"https:\/\/deeplearningindaba.com\/2026\/accepted-african-datasets\/","title":{"rendered":"Accepted African Datasets"},"content":{"rendered":"<div class=\"et_pb_section_0 et_pb_section et_pb_fullwidth_section et_section_regular et_block_section\"><span class=\"et_pb_background_pattern\"><\/span>\n<section class=\"et_pb_fullwidth_header_0 et_pb_fullwidth_header et_pb_bg_layout_dark et_pb_text_align_left et_pb_module et_flex_module\"><div class=\"et_pb_fullwidth_header_container left\"><div class=\"header-content-container center\"><div class=\"header-content et_flex_module\"><h1 class=\"et_pb_module_header\"> Accepted African Datasets<\/h1><div class=\"et_pb_header_button_wrapper\"><\/div><\/div><\/div><\/div><div class=\"et_pb_fullwidth_header_overlay\"><\/div><div class=\"et_pb_fullwidth_header_scroll\"><\/div><\/section>\n<\/div>\n\n<div class=\"et_pb_section_1 et_pb_section et_section_regular et_flex_section preset--module--divi-section--default\">\n<div class=\"et_pb_row_0 et_pb_row et_flex_row\">\n<div class=\"et_pb_column_0 et_pb_column et-last-child et_flex_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tablet et_flex_column_24_24_phone\">\n<div class=\"et_pb_heading_0 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--c33f07d9-41e0-421a-8799-5799df695cce\"><div class=\"et_pb_heading_container\"><h2 class=\"et_pb_module_header\">Accepted Datasets\u00a0<\/h2><\/div><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_1 et_pb_row et_pb_row_3-4_1-4 et_block_row et_block_row_3-4_1-4\">\n<div class=\"et_pb_column_1 et_pb_column et_pb_column_3_4 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_1 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--ba4a6336-701f-47b8-bf5c-09da0ce28016\"><div class=\"et_pb_heading_container\"><h3 class=\"et_pb_module_header\">Chichewa Text-to-SQL Benchmark Dataset<\/h3><\/div><\/div>\n\n<div class=\"et_pb_text_0 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--2c55a9c4-feed-423b-9edb-ae0b5b365cac\"><div class=\"et_pb_text_inner\"><p><span data-sheets-root=\"1\"><strong>Presenter Name<\/strong>: John Emeka Eze<br \/><strong>Dataset Domain<\/strong>: Agriculture, Education, Language &amp; Culture | Langue et culture, Infrastructure<br \/><strong>Video Presentation<\/strong>: <\/span><span data-sheets-root=\"1\"><a href=\"https:\/\/drive.google.com\/open?id=18Mp1HnlZg8Rp4yuzAIle0XTDRcSAr-gb\">https:\/\/drive.google.com\/open?id=18Mp1HnlZg8Rp4yuzAIle0XTDRcSAr-gb<\/a>\u00a0<\/span><\/p>\n<p><span data-sheets-root=\"1\"><strong>Dataset Description:<\/strong> The Chichewa Text-to-SQL dataset is a curated benchmark designed to support research in natural language interfaces for low-resource African languages. It consists of 400 natural language\u2013SQL query pairs, where each query is written in Chichewa and mapped to structured SQL over a unified relational database. The dataset covers practical domains such as agriculture, commodity prices, population statistics, and food security, reflecting real-world information needs in Malawi and similar contexts. It was developed through careful translation and validation to ensure semantic consistency between queries and database schemas. This dataset enables the evaluation and development of machine learning models, particularly large language models, for structured query generation in data-constrained and linguistically underrepresented environments.<br \/><\/span><\/p>\n<\/div><\/div>\n<\/div>\n\n<div class=\"et_pb_column_2 et_pb_column et_pb_column_1_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_2 et_pb_heading et_pb_module et_block_module ai_ignore_all preset--module--divi-heading--e8a6df2b-53a6-4b58-8bda-5426d033ba1a\"><div class=\"et_pb_heading_container\"><h6 class=\"et_pb_module_header\">Nigeria<\/h6><\/div><\/div>\n\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_0_wrapper preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206_wrapper\"><a class=\"et_pb_button_0 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206\" href=\"https:\/\/huggingface.co\/datasets\/johneze\/chichewa-text2sql\" target=\"_blank\">Learn More<\/a><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_2 et_pb_row et_pb_row_3-4_1-4 et_block_row et_block_row_3-4_1-4\">\n<div class=\"et_pb_column_3 et_pb_column et_pb_column_3_4 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_3 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--ba4a6336-701f-47b8-bf5c-09da0ce28016\"><div class=\"et_pb_heading_container\"><h3 class=\"et_pb_module_header\">Cactus_dataset_V2<\/h3><\/div><\/div>\n\n<div class=\"et_pb_text_1 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--2c55a9c4-feed-423b-9edb-ae0b5b365cac\"><div class=\"et_pb_text_inner\"><p><span data-sheets-root=\"1\"><strong>Presenter Name<\/strong>: Adel BENALI<br \/><strong>Dataset Domain<\/strong>: Agriculture<br \/><strong>Citation: <\/strong>BENALI, A., Fourati, R., &amp; JDEY, I. (2026). Cactus Disease dataset V2 (Version 2) [Data set]. Zenodo.<\/span><\/p>\n<p><span data-sheets-root=\"1\"><strong>Dataset Description:<\/strong> This repository contains two image datasets of cacti designed for scientific research in plant disease classification and segmentation using machine learning. All images are labeled and organized in folders according to their respective classes.<\/p>\n<p><\/span><\/p>\n<\/div><\/div>\n<\/div>\n\n<div class=\"et_pb_column_4 et_pb_column et_pb_column_1_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_4 et_pb_heading et_pb_module et_block_module ai_ignore_all preset--module--divi-heading--e8a6df2b-53a6-4b58-8bda-5426d033ba1a\"><div class=\"et_pb_heading_container\"><h6 class=\"et_pb_module_header\">Tunisia<\/h6><\/div><\/div>\n\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_1_wrapper preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206_wrapper\"><a class=\"et_pb_button_1 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206\" href=\"https:\/\/doi.org\/10.5281\/zenodo.19075056\" target=\"_blank\">Learn More<\/a><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_3 et_pb_row et_pb_row_3-4_1-4 et_block_row et_block_row_3-4_1-4\">\n<div class=\"et_pb_column_5 et_pb_column et_pb_column_3_4 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_5 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--ba4a6336-701f-47b8-bf5c-09da0ce28016\"><div class=\"et_pb_heading_container\"><h3 class=\"et_pb_module_header\">OHADA-CCJA Court Decisions Corpus: Case Law of the Common Court of Justice and Arbitration for African Legal NLP\u00a0<\/h3><\/div><\/div>\n\n<div class=\"et_pb_text_2 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--2c55a9c4-feed-423b-9edb-ae0b5b365cac\"><div class=\"et_pb_text_inner\"><p><span data-sheets-root=\"1\"><strong>Presenter Name<\/strong>: Foutse Yuehgoh<br \/><strong>Dataset Domain<\/strong>: <span style=\"font-weight: 400;\">Legal &amp; Economic Development \/ Law &amp; Governance<\/span><br \/><strong>Video Presentation<\/strong>: <a href=\"https:\/\/drive.google.com\/open?id=1eSSDrmBpEK7UPnN7xxD1WNqvVvQ8a6wz\">https:\/\/drive.google.com\/open?id=1eSSDrmBpEK7UPnN7xxD1WNqvVvQ8a6wz<\/a> <\/span><span data-sheets-root=\"1\">\u00a0<\/span><\/p>\n<p><span data-sheets-root=\"1\"><strong>Dataset Description:<\/strong> A curated corpus of 4,059 court decisions from the Cour Commune de Justice et d'Arbitrage (CCJA), the supranational court of the Organisation pour l'Harmonisation en Afrique du Droit des Affaires (OHADA). OHADA harmonizes business law across 17 African member states: Benin, Burkina Faso, Cameroon, Central African Republic, Chad, Comoros, Democratic Republic of Congo, Republic of Congo, C\u00f4te d'Ivoire, Equatorial Guinea, Gabon, Guinea, Guinea-Bissau, Mali, Niger, Senegal, and Togo. This dataset provides structured access to CCJA jurisprudence spanning over two decades (1997\u20132023), making it a unique resource for African legal NLP research.<br \/><\/span><\/p>\n<\/div><\/div>\n<\/div>\n\n<div class=\"et_pb_column_6 et_pb_column et_pb_column_1_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_6 et_pb_heading et_pb_module et_block_module ai_ignore_all preset--module--divi-heading--e8a6df2b-53a6-4b58-8bda-5426d033ba1a\"><div class=\"et_pb_heading_container\"><h6 class=\"et_pb_module_header\">France<\/h6><\/div><\/div>\n\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_2_wrapper preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206_wrapper\"><a class=\"et_pb_button_2 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206\" href=\"https:\/\/huggingface.co\/datasets\/johneze\/chichewa-text2sql\" target=\"_blank\">Learn More<\/a><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_4 et_pb_row et_pb_row_3-4_1-4 et_block_row et_block_row_3-4_1-4\">\n<div class=\"et_pb_column_7 et_pb_column et_pb_column_3_4 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_7 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--ba4a6336-701f-47b8-bf5c-09da0ce28016\"><div class=\"et_pb_heading_container\"><h3 class=\"et_pb_module_header\">UCCB: Uganda Cultural Context Benchmark<\/h3><\/div><\/div>\n\n<div class=\"et_pb_text_3 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--2c55a9c4-feed-423b-9edb-ae0b5b365cac\"><div class=\"et_pb_text_inner\"><p><span data-sheets-root=\"1\"><strong>Presenter Name<\/strong>: Lameck Kavuma<br \/><strong>Dataset Domain<\/strong>: <span style=\"font-weight: 400;\">Language &amp; Culture | Langue et culture<\/span><br \/><strong>Video Presentation<\/strong>: <a href=\"https:\/\/drive.google.com\/drive\/folders\/1gIMKyaN27CgpGl9hecoiF3CsFTgQVd-X?usp=sharing\">https:\/\/drive.google.com\/drive\/folders\/1gIMKyaN27CgpGl9hecoiF3CsFTgQVd-X?usp=sharing<\/a>\u00a0<\/span><\/p>\n<p><span data-sheets-root=\"1\"><strong>Full data card<\/strong>: <a href=\"https:\/\/huggingface.co\/datasets\/CraneAILabs\/UCCB\">https:\/\/huggingface.co\/datasets\/CraneAILabs\/UCCB<\/a>\u00a0<br \/><strong>GitHub repository:<\/strong> <a href=\"https:\/\/github.com\/Crane-AI-Labs\/UCCB\">https:\/\/github.com\/Crane-AI-Labs\/UCCB<\/a>\u00a0<br \/><strong>UK AI Safety Institute Inspect AI integration:<\/strong> <a href=\"https:\/\/github.com\/UKGovernmentBEIS\/inspect_evals\/tree\/main\/src\/inspect_evals\/uccb\">https:\/\/github.com\/UKGovernmentBEIS\/inspect_evals\/tree\/main\/src\/inspect_evals\/uccb<\/a>\u00a0<br \/><\/span><\/p>\n<p><span data-sheets-root=\"1\"><strong>Dataset Description:<\/strong> The Uganda Cultural Context Benchmark (UCCB) is the first comprehensive question-answer dataset designed to evaluate the cultural understanding and reasoning abilities of Large Language Models (LLMs) concerning Uganda. It contains 1,039 expert-validated question-answer pairs across 24 cultural domains including folklore, traditional medicine, music, attire, slang, history, religion, and economy. The benchmark incorporates terminology from multiple indigenous Ugandan languages (Luganda, Runyankole, Acholi, Ateso, Lugbara, among others) embedded within English-language questions. UCCB was curated through a hybrid pipeline combining Wikipedia-sourced content with LLM-assisted generation, chunk-grounded quality scoring, and validation by 10 Ugandan annotators from across the country's four major regions. The dataset has been integrated into the UK AI Safety Institute's Inspect AI framework, marking the first African cultural benchmark in a sovereign government's AI evaluation infrastructure. Evaluation of five frontier LLMs revealed a 2.94-point spread on a 5-point scale, demonstrating significant variance in cultural knowledge across model families.<\/span><\/p>\n<p>&nbsp;<\/p>\n<\/div><\/div>\n<\/div>\n\n<div class=\"et_pb_column_8 et_pb_column et_pb_column_1_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_8 et_pb_heading et_pb_module et_block_module ai_ignore_all preset--module--divi-heading--e8a6df2b-53a6-4b58-8bda-5426d033ba1a\"><div class=\"et_pb_heading_container\"><h6 class=\"et_pb_module_header\">Uganda<\/h6><\/div><\/div>\n\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_3_wrapper preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206_wrapper\"><a class=\"et_pb_button_3 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206\" href=\"https:\/\/huggingface.co\/datasets\/CraneAILabs\/UCCB\" target=\"_blank\">Learn More<\/a><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_5 et_pb_row et_pb_row_3-4_1-4 et_block_row et_block_row_3-4_1-4\">\n<div class=\"et_pb_column_9 et_pb_column et_pb_column_3_4 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_9 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--ba4a6336-701f-47b8-bf5c-09da0ce28016\"><div class=\"et_pb_heading_container\"><h3 class=\"et_pb_module_header\">African Urban Land Use and Land Cover Segmentation Dataset (AULC-13)<\/h3><\/div><\/div>\n\n<div class=\"et_pb_text_4 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--2c55a9c4-feed-423b-9edb-ae0b5b365cac\"><div class=\"et_pb_text_inner\"><p><span data-sheets-root=\"1\"><strong>Presenter Name<\/strong>: Benayad Mohamed<br \/><strong>Dataset Domain<\/strong>: Agriculture, Education, Environment | Environnement, Infrastructure<\/span><span data-sheets-root=\"1\">\u00a0<br \/><\/span><\/p>\n<p><span data-sheets-root=\"1\"><strong>Dataset Description:<\/strong><\/span><\/p>\n<p>AfriUrban-13 is a high-resolution satellite image dataset designed for semantic segmentation of urban land use and land cover in African cities. The dataset contains manually annotated RGB satellite image tiles labeled into 13 classes: roads, vegetation, grass, trees, brown soil, bare land, buildings, solar panels, tracks, sidewalks, playgrounds, cars, and water. The dataset was developed to address the underrepresentation of African urban environments in existing global segmentation benchmarks. It captures the spatial complexity of semi-arid urban areas, including mixed land cover patterns, informal structures, distributed solar installations, and transportation infrastructure. AfriUrban-13 is intended to support the development and evaluation of deep learning models, including CNN-based architectures (e.g., U-Net) and transformer-based segmentation models (e.g., SegFormer). It can be used for applications in smart city planning, infrastructure development, renewable energy mapping, environmental monitoring, and sustainable urban development. By providing high-quality pixel-level annotations from an African context, this dataset aims to strengthen locally relevant AI research and contribute to more inclusive global geospatial benchmarks.<\/p>\n<p><span style=\"font-weight: 400;\">\u00a0License: Creative Commons Attribution 4.0 International (CC BY 4.0)\u00a0<\/span><\/p>\n<p>&nbsp;<\/p>\n<\/div><\/div>\n<\/div>\n\n<div class=\"et_pb_column_10 et_pb_column et_pb_column_1_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_10 et_pb_heading et_pb_module et_block_module ai_ignore_all preset--module--divi-heading--e8a6df2b-53a6-4b58-8bda-5426d033ba1a\"><div class=\"et_pb_heading_container\"><h6 class=\"et_pb_module_header\">Morocco<\/h6><\/div><\/div>\n\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_4_wrapper preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206_wrapper\"><a class=\"et_pb_button_4 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206\" href=\"https:\/\/drive.google.com\/drive\/folders\/1EHq_3QhHzYJwF3KsVbrfcKSfVov65cGx?usp=drive_link\" target=\"_blank\">Learn More<\/a><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_6 et_pb_row et_pb_row_3-4_1-4 et_block_row et_block_row_3-4_1-4\">\n<div class=\"et_pb_column_11 et_pb_column et_pb_column_3_4 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_11 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--ba4a6336-701f-47b8-bf5c-09da0ce28016\"><div class=\"et_pb_heading_container\"><h3 class=\"et_pb_module_header\">AI Startups in Africa<\/h3><\/div><\/div>\n\n<div class=\"et_pb_text_5 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--2c55a9c4-feed-423b-9edb-ae0b5b365cac\"><div class=\"et_pb_text_inner\"><p><span data-sheets-root=\"1\"><strong>Presenter Name<\/strong>: Chinasa T. Okolo<br \/><strong>Dataset Domain<\/strong>: Industry<\/span><span data-sheets-root=\"1\">\u00a0<br \/><\/span><\/p>\n<p><span data-sheets-root=\"1\"><strong>Video Presentation<\/strong>: <a href=\"https:\/\/drive.google.com\/open?id=1AxSR4zUatstUbU3pTfUNyKyG3idQfMNp\">https:\/\/drive.google.com\/open?id=1AxSR4zUatstUbU3pTfUNyKyG3idQfMNp<\/a>\u00a0\u00a0<\/span><\/p>\n<p><span data-sheets-root=\"1\"><strong>Dataset Description: <\/strong><\/span>A list of AI startups in Africa with company name, sector\/domain, location, customer type, year founded, team, stage, operating status, and brief description.<\/p>\n<\/div><\/div>\n<\/div>\n\n<div class=\"et_pb_column_12 et_pb_column et_pb_column_1_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_12 et_pb_heading et_pb_module et_block_module ai_ignore_all preset--module--divi-heading--e8a6df2b-53a6-4b58-8bda-5426d033ba1a\"><div class=\"et_pb_heading_container\"><h6 class=\"et_pb_module_header\">United States<\/h6><\/div><\/div>\n\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_5_wrapper preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206_wrapper\"><a class=\"et_pb_button_5 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206\" href=\"https:\/\/dataverse.harvard.edu\/dataset.xhtml?persistentId=doi:10.7910\/DVN\/Z0PKGO&#038;version=1.0\" target=\"_blank\">Learn More<\/a><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_7 et_pb_row et_pb_row_3-4_1-4 et_block_row et_block_row_3-4_1-4\">\n<div class=\"et_pb_column_13 et_pb_column et_pb_column_3_4 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_13 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--ba4a6336-701f-47b8-bf5c-09da0ce28016\"><div class=\"et_pb_heading_container\"><h3 class=\"et_pb_module_header\">Senegal Maternal Health QA dataset (SenMH-QA)<\/h3><\/div><\/div>\n\n<div class=\"et_pb_text_6 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--2c55a9c4-feed-423b-9edb-ae0b5b365cac\"><div class=\"et_pb_text_inner\"><p><span data-sheets-root=\"1\"><strong>Presenter Name<\/strong>: Ertony Basilwango<br \/><strong>Dataset Domain<\/strong>: <span style=\"font-weight: 400;\">Health | Sant\u00e9, Language &amp; Culture | Langue et culture<\/span><br \/><strong>Video Presentation<\/strong>: <a href=\"https:\/\/drive.google.com\/open?id=1ATVDQUqCR0eGbJDqHC9UbZ3s-0aV4eLg\">https:\/\/drive.google.com\/open?id=1ATVDQUqCR0eGbJDqHC9UbZ3s-0aV4eLg<\/a>\u00a0<\/span><\/p>\n<p><span data-sheets-root=\"1\"><a href=\"https:\/\/drive.google.com\/open?id=1eSSDrmBpEK7UPnN7xxD1WNqvVvQ8a6wz\"><\/a><\/span><span data-sheets-root=\"1\"><strong>Dataset Description:<\/strong> This dataset contains short audio recordings collected in community settings in Senegal. The recordings cover maternal health topics, including family planning, healthcare access, pregnancy practices, and cultural beliefs around maternal and reproductive health. Audio was captured using mobile devices or portable recorders in natural conversational conditions, and all transcriptions were manually verified by Wolof linguists, maternal health experts, and machine learning engineers. The dataset is intended for training and benchmarking domain-specific automatic speech recognition (ASR) models.<\/p>\n<p><\/span><\/p>\n<\/div><\/div>\n<\/div>\n\n<div class=\"et_pb_column_14 et_pb_column et_pb_column_1_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_14 et_pb_heading et_pb_module et_block_module ai_ignore_all preset--module--divi-heading--e8a6df2b-53a6-4b58-8bda-5426d033ba1a\"><div class=\"et_pb_heading_container\"><h6 class=\"et_pb_module_header\">Senegal<\/h6><\/div><\/div>\n\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_6_wrapper preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206_wrapper\"><a class=\"et_pb_button_6 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206\" href=\"\" target=\"_blank\">Learn More<\/a><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_8 et_pb_row et_pb_row_3-4_1-4 et_block_row et_block_row_3-4_1-4\">\n<div class=\"et_pb_column_15 et_pb_column et_pb_column_3_4 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_15 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--ba4a6336-701f-47b8-bf5c-09da0ce28016\"><div class=\"et_pb_heading_container\"><h3 class=\"et_pb_module_header\">AfricaBias-SW-FR: A Multilingual Bias Detection Dataset for Swahili and French<\/h3><\/div><\/div>\n\n<div class=\"et_pb_text_7 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--2c55a9c4-feed-423b-9edb-ae0b5b365cac\"><div class=\"et_pb_text_inner\"><p><span data-sheets-root=\"1\"><strong>Presenter Name<\/strong>: Preston Osoro<br \/><strong>Dataset Domain<\/strong>: Language &amp; Culture | Langue et culture<br \/><\/span><\/p>\n<p><span data-sheets-root=\"1\"><a href=\"https:\/\/drive.google.com\/open?id=1eSSDrmBpEK7UPnN7xxD1WNqvVvQ8a6wz\"><\/a><\/span><span data-sheets-root=\"1\"><strong>Dataset Description:<\/strong> AfricaBias-SW-FR is a multilingual bias detection dataset containing 35,285 annotated sentences in Swahili and French, developed to address the near-total absence of bias benchmarks for African languages. Each sentence is labelled across four categories, neutral, stereotype, counter-stereotype, and derogation, spanning eight social domains, including household and care, livelihoods and work, governance, and health. The dataset draws from real Kenyan and Francophone African sources, including news, social media, encyclopedias, and government documents, and captures gender, religious, and cultural sensitivity attributes. It is the first publicly released large-scale gender and social bias detection dataset for Swahili, produced as a Gates Foundation-supported deliverable through the AfriLabs AI Accelerator programme.<br \/><\/span><\/p>\n<p>&nbsp;<\/p>\n<\/div><\/div>\n<\/div>\n\n<div class=\"et_pb_column_16 et_pb_column et_pb_column_1_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_16 et_pb_heading et_pb_module et_block_module ai_ignore_all preset--module--divi-heading--e8a6df2b-53a6-4b58-8bda-5426d033ba1a\"><div class=\"et_pb_heading_container\"><h6 class=\"et_pb_module_header\">Kenya<\/h6><\/div><\/div>\n\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_7_wrapper preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206_wrapper\"><a class=\"et_pb_button_7 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206\" href=\"https:\/\/huggingface.co\/datasets\/Daudipdg\/africabias-sw-fr\" target=\"_blank\">Learn More<\/a><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_9 et_pb_row et_pb_row_3-4_1-4 et_block_row et_block_row_3-4_1-4\">\n<div class=\"et_pb_column_17 et_pb_column et_pb_column_3_4 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_17 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--ba4a6336-701f-47b8-bf5c-09da0ce28016\"><div class=\"et_pb_heading_container\"><h3 class=\"et_pb_module_header\">mghana_st<\/h3><\/div><\/div>\n\n<div class=\"et_pb_text_8 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--2c55a9c4-feed-423b-9edb-ae0b5b365cac\"><div class=\"et_pb_text_inner\"><p><span data-sheets-root=\"1\"><strong>Presenter Name<\/strong>: Jesse Johnson<br \/><strong>Dataset Domain<\/strong>: Language &amp; Culture | Langue et culture<br \/><\/span><\/p>\n<p><span data-sheets-root=\"1\"><a href=\"https:\/\/drive.google.com\/open?id=1eSSDrmBpEK7UPnN7xxD1WNqvVvQ8a6wz\"><\/a><\/span><span data-sheets-root=\"1\"><strong>Dataset Description:<\/strong> mghana-st is a curated audio dataset intended for speech tasks such as: speech recognition, speaker identification, speech clustering, or speech translation.<br \/><\/span><\/p>\n<p>&nbsp;<\/p>\n<\/div><\/div>\n<\/div>\n\n<div class=\"et_pb_column_18 et_pb_column et_pb_column_1_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_18 et_pb_heading et_pb_module et_block_module ai_ignore_all preset--module--divi-heading--e8a6df2b-53a6-4b58-8bda-5426d033ba1a\"><div class=\"et_pb_heading_container\"><h6 class=\"et_pb_module_header\">Ghana<\/h6><\/div><\/div>\n\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_8_wrapper preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206_wrapper\"><a class=\"et_pb_button_8 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206\" href=\"https:\/\/huggingface.co\/datasets\/adwumatech-ai\/mghana-st\" target=\"_blank\">Learn More<\/a><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_10 et_pb_row et_pb_row_3-4_1-4 et_block_row et_block_row_3-4_1-4\">\n<div class=\"et_pb_column_19 et_pb_column et_pb_column_3_4 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_19 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--ba4a6336-701f-47b8-bf5c-09da0ce28016\"><div class=\"et_pb_heading_container\"><h3 class=\"et_pb_module_header\">Enugu Dialect Igbo language and cultural multimodal dataset<\/h3><\/div><\/div>\n\n<div class=\"et_pb_text_9 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--2c55a9c4-feed-423b-9edb-ae0b5b365cac\"><div class=\"et_pb_text_inner\"><p><span data-sheets-root=\"1\"><strong>Presenter Name<\/strong>: Peter Emmanuel Olotuche<br \/><strong>Dataset Domain<\/strong>: <span style=\"font-weight: 400;\">Education, Environment | Environnement, Language &amp; Culture | Langue et culture<\/span><br \/><\/span><\/p>\n<p><span data-sheets-root=\"1\"><a href=\"https:\/\/drive.google.com\/open?id=1eSSDrmBpEK7UPnN7xxD1WNqvVvQ8a6wz\"><\/a><\/span><span data-sheets-root=\"1\"><strong>Dataset Description:<\/strong> This dataset contains multimodal data representing the Igbo language and culture from Enugu State, Nigeria. It includes text documents, images, audio recordings, and short videos capturing everyday conversations, cultural expressions, environmental contexts, and educational materials. The dataset is designed to support AI training, language preservation, and research in African languages and cultural knowledge.<br \/><\/span><\/p>\n<p>&nbsp;<\/p>\n<\/div><\/div>\n<\/div>\n\n<div class=\"et_pb_column_20 et_pb_column et_pb_column_1_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_20 et_pb_heading et_pb_module et_block_module ai_ignore_all preset--module--divi-heading--e8a6df2b-53a6-4b58-8bda-5426d033ba1a\"><div class=\"et_pb_heading_container\"><h6 class=\"et_pb_module_header\">Nigeria<\/h6><\/div><\/div>\n\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_9_wrapper preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206_wrapper\"><a class=\"et_pb_button_9 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206\" href=\"https:\/\/drive.google.com\/file\/d\/1W9zHbFGgQ4FQIVqhbuYNhxnZcmx39R0c\/view?usp=drivesdk\" target=\"_blank\">Learn More<\/a><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_11 et_pb_row et_pb_row_3-4_1-4 et_block_row et_block_row_3-4_1-4\">\n<div class=\"et_pb_column_21 et_pb_column et_pb_column_3_4 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_21 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--ba4a6336-701f-47b8-bf5c-09da0ce28016\"><div class=\"et_pb_heading_container\"><h3 class=\"et_pb_module_header\">African Plums Dataset<\/h3><\/div><\/div>\n\n<div class=\"et_pb_text_10 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--2c55a9c4-feed-423b-9edb-ae0b5b365cac\"><div class=\"et_pb_text_inner\"><p><span data-sheets-root=\"1\"><strong>Presenter Name<\/strong>: FAMENI TAGNI ARMEL GABIN<br \/><strong>Dataset Domain<\/strong>: Agriculture<br \/><\/span><\/p>\n<p><span data-sheets-root=\"1\"><a href=\"https:\/\/drive.google.com\/open?id=1eSSDrmBpEK7UPnN7xxD1WNqvVvQ8a6wz\"><\/a><\/span><span data-sheets-root=\"1\"><strong>Dataset Description:<\/strong> This dataset contains 4,507 annotated images of African plums collected from various fields in Cameroon. It is specifically designed for training and evaluating AI models in fruit quality assessment and defect detection. The images are categorized into six classes based on their defect type: bruised, cracked, rotten, spotted, unaffected, and unripe.<\/span><\/p>\n<\/div><\/div>\n<\/div>\n\n<div class=\"et_pb_column_22 et_pb_column et_pb_column_1_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_22 et_pb_heading et_pb_module et_block_module ai_ignore_all preset--module--divi-heading--e8a6df2b-53a6-4b58-8bda-5426d033ba1a\"><div class=\"et_pb_heading_container\"><h6 class=\"et_pb_module_header\">Cameroon<\/h6><\/div><\/div>\n\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_10_wrapper preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206_wrapper\"><a class=\"et_pb_button_10 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206\" href=\"https:\/\/www.kaggle.com\/dsv\/9694239\" target=\"_blank\">Learn More<\/a><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_12 et_pb_row et_pb_row_3-4_1-4 et_block_row et_block_row_3-4_1-4\">\n<div class=\"et_pb_column_23 et_pb_column et_pb_column_3_4 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_23 et_pb_heading et_pb_module et_block_module preset--module--divi-heading--ba4a6336-701f-47b8-bf5c-09da0ce28016\"><div class=\"et_pb_heading_container\"><h3 class=\"et_pb_module_header\">FITSIPIKA Malagasy Dataset (SOKAJY)<\/h3><\/div><\/div>\n\n<div class=\"et_pb_text_11 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--2c55a9c4-feed-423b-9edb-ae0b5b365cac\"><div class=\"et_pb_text_inner\"><p><span data-sheets-root=\"1\"><strong>Presenter Name<\/strong>: <span style=\"font-weight: 400;\">Vatosoa Razafindrazaka<\/span><br \/><strong>Dataset Domain<\/strong>: Language &amp; Culture | Langue et culture<br \/><\/span><\/p>\n<p><span data-sheets-root=\"1\"><a href=\"https:\/\/drive.google.com\/open?id=1eSSDrmBpEK7UPnN7xxD1WNqvVvQ8a6wz\"><\/a><\/span><span data-sheets-root=\"1\"><strong>Dataset Description:<\/strong> A Part-of-speech corpus for Malagasy (28M Speakers) created to bridge the gap in low-resource NLP. It includes manual annotations for 4,693 sentences. Validated through peer-reviewed publication in Springer Nature, this dataset serves as a foundation for building ASR Malagasy and transcription systems.<br \/><\/span><\/p>\n<p><span data-sheets-root=\"1\">License: CC BY-NC-SA 4.0\u00a0<br \/><\/span><\/p>\n<p>&nbsp;<\/p>\n<\/div><\/div>\n<\/div>\n\n<div class=\"et_pb_column_24 et_pb_column et_pb_column_1_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_heading_24 et_pb_heading et_pb_module et_block_module ai_ignore_all preset--module--divi-heading--e8a6df2b-53a6-4b58-8bda-5426d033ba1a\"><div class=\"et_pb_heading_container\"><h6 class=\"et_pb_module_header\">Madagascar <\/h6><\/div><\/div>\n\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_11_wrapper preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206_wrapper\"><a class=\"et_pb_button_11 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-button--1bffc0fc-42a0-49d1-bd9a-ae3ade2d7206\" href=\"https:\/\/huggingface.co\/datasets\/Vatosoa\/pos-tagging-malagasy-sokajy\" target=\"_blank\">Learn More<\/a><\/div>\n<\/div>\n<\/div>\n<\/div>","protected":false},"excerpt":{"rendered":"","protected":false},"author":32,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"_acf_changed":false,"footnotes":"","_links_to":"","_links_to_target":""},"class_list":["post-1214","page","type-page","status-publish","hentry"],"acf":[],"_links":{"self":[{"href":"https:\/\/deeplearningindaba.com\/2026\/wp-json\/wp\/v2\/pages\/1214","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/deeplearningindaba.com\/2026\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/deeplearningindaba.com\/2026\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/deeplearningindaba.com\/2026\/wp-json\/wp\/v2\/users\/32"}],"replies":[{"embeddable":true,"href":"https:\/\/deeplearningindaba.com\/2026\/wp-json\/wp\/v2\/comments?post=1214"}],"version-history":[{"count":10,"href":"https:\/\/deeplearningindaba.com\/2026\/wp-json\/wp\/v2\/pages\/1214\/revisions"}],"predecessor-version":[{"id":1259,"href":"https:\/\/deeplearningindaba.com\/2026\/wp-json\/wp\/v2\/pages\/1214\/revisions\/1259"}],"wp:attachment":[{"href":"https:\/\/deeplearningindaba.com\/2026\/wp-json\/wp\/v2\/media?parent=1214"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}