ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data

Despite significant research on the Bangla language in Natural Language Processing (NLP), there remains a notable resource deficit for its diverse regional dialects, such as those spoken in Chittagong, Sylhet, and Barisal. These dialects, often considered unintelligible to speakers of Standard Benga...

Full description

Saved in:
Bibliographic Details
Main Authors: Nusrat Sultana, Rumana Yasmin, Bijon Mallik, Mohammad Shorif Uddin
Format: Article
Language:English
Published: Elsevier 2025-02-01
Series:Data in Brief
Subjects:
Online Access:http://www.sciencedirect.com/science/article/pii/S2352340925000083
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1832576508927934464
author Nusrat Sultana
Rumana Yasmin
Bijon Mallik
Mohammad Shorif Uddin
author_facet Nusrat Sultana
Rumana Yasmin
Bijon Mallik
Mohammad Shorif Uddin
author_sort Nusrat Sultana
collection DOAJ
description Despite significant research on the Bangla language in Natural Language Processing (NLP), there remains a notable resource deficit for its diverse regional dialects, such as those spoken in Chittagong, Sylhet, and Barisal. These dialects, often considered unintelligible to speakers of Standard Bengali, pose challenges due to their unique grammatical structures and phonetic variations. Some linguists categorize them as distinct languages. To address this, we present ONUBAD, a large and freely available dataset for the automatic translation of Chittagong, Sylhet, and Barisal dialects into Standard Bangla using a Neural Machine Translation (NMT) system. ONUBAD provides a parallel corpus of 1540 words, 130 clauses, and 980 sentences per regional dialect and their standard counterparts along with English translation. The dataset includes metadata on phonetic variations and grammatical features, aiming to bridge the gap between standard and non-standard forms of Bangla. It serves as a valuable resource for researchers in NLP, dialect studies, and linguistic preservation, helping to develop more accurate and contextually relevant translation models. The dataset was collected between July and September 2024 from diverse sources such as books, websites, and regional people with the help of regional dialect specialists. It is hosted by the Department of Computer Science and Engineering, Jahangirnagar University, and is freely accessible at https://data.mendeley.com/datasets/6ft99kf89b/2
format Article
id doaj-art-d1ffad4a0c6343a28dc8c705505d7fb6
institution Kabale University
issn 2352-3409
language English
publishDate 2025-02-01
publisher Elsevier
record_format Article
series Data in Brief
spelling doaj-art-d1ffad4a0c6343a28dc8c705505d7fb62025-01-31T05:11:47ZengElsevierData in Brief2352-34092025-02-0158111276ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley DataNusrat Sultana0Rumana Yasmin1Bijon Mallik2Mohammad Shorif Uddin3Department of Computer Science and Engineering, Jahangirnagar University, Dhaka, Bangladesh; Department of Computer Science and Engineering, Bangladesh University of Business and Technology, Dhaka, Bangladesh; Corresponding author at: Department of Computer Science and Engineering, Jahangirnagar University, Dhaka, Bangladesh.Department of Computer Science and Engineering, Jahangirnagar University, Dhaka, Bangladesh; Department of Computer Science and Engineering, Bangladesh University of Professionals, Dhaka, BangladeshDepartment of Computer Science and Engineering, Bangladesh University of Engineering and Technology, Dhaka, Bangladesh; Department of Computer Science and Engineering, Bangladesh University of Business and Technology, Dhaka, BangladeshDepartment of Computer Science and Engineering, Jahangirnagar University, Dhaka, BangladeshDespite significant research on the Bangla language in Natural Language Processing (NLP), there remains a notable resource deficit for its diverse regional dialects, such as those spoken in Chittagong, Sylhet, and Barisal. These dialects, often considered unintelligible to speakers of Standard Bengali, pose challenges due to their unique grammatical structures and phonetic variations. Some linguists categorize them as distinct languages. To address this, we present ONUBAD, a large and freely available dataset for the automatic translation of Chittagong, Sylhet, and Barisal dialects into Standard Bangla using a Neural Machine Translation (NMT) system. ONUBAD provides a parallel corpus of 1540 words, 130 clauses, and 980 sentences per regional dialect and their standard counterparts along with English translation. The dataset includes metadata on phonetic variations and grammatical features, aiming to bridge the gap between standard and non-standard forms of Bangla. It serves as a valuable resource for researchers in NLP, dialect studies, and linguistic preservation, helping to develop more accurate and contextually relevant translation models. The dataset was collected between July and September 2024 from diverse sources such as books, websites, and regional people with the help of regional dialect specialists. It is hosted by the Department of Computer Science and Engineering, Jahangirnagar University, and is freely accessible at https://data.mendeley.com/datasets/6ft99kf89b/2http://www.sciencedirect.com/science/article/pii/S2352340925000083Neural machine translationNatural language processingRegional language datasetText translationStandard Bangla language
spellingShingle Nusrat Sultana
Rumana Yasmin
Bijon Mallik
Mohammad Shorif Uddin
ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data
Data in Brief
Neural machine translation
Natural language processing
Regional language dataset
Text translation
Standard Bangla language
title ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data
title_full ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data
title_fullStr ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data
title_full_unstemmed ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data
title_short ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data
title_sort onubad a comprehensive dataset for automated conversion of bangla regional dialects into standard bengali dialectmendeley data
topic Neural machine translation
Natural language processing
Regional language dataset
Text translation
Standard Bangla language
url http://www.sciencedirect.com/science/article/pii/S2352340925000083
work_keys_str_mv AT nusratsultana onubadacomprehensivedatasetforautomatedconversionofbanglaregionaldialectsintostandardbengalidialectmendeleydata
AT rumanayasmin onubadacomprehensivedatasetforautomatedconversionofbanglaregionaldialectsintostandardbengalidialectmendeleydata
AT bijonmallik onubadacomprehensivedatasetforautomatedconversionofbanglaregionaldialectsintostandardbengalidialectmendeleydata
AT mohammadshorifuddin onubadacomprehensivedatasetforautomatedconversionofbanglaregionaldialectsintostandardbengalidialectmendeleydata