ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data
Despite significant research on the Bangla language in Natural Language Processing (NLP), there remains a notable resource deficit for its diverse regional dialects, such as those spoken in Chittagong, Sylhet, and Barisal. These dialects, often considered unintelligible to speakers of Standard Benga...
Saved in:
Main Authors: | , , , |
---|---|
Format: | Article |
Language: | English |
Published: |
Elsevier
2025-02-01
|
Series: | Data in Brief |
Subjects: | |
Online Access: | http://www.sciencedirect.com/science/article/pii/S2352340925000083 |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
_version_ | 1832576508927934464 |
---|---|
author | Nusrat Sultana Rumana Yasmin Bijon Mallik Mohammad Shorif Uddin |
author_facet | Nusrat Sultana Rumana Yasmin Bijon Mallik Mohammad Shorif Uddin |
author_sort | Nusrat Sultana |
collection | DOAJ |
description | Despite significant research on the Bangla language in Natural Language Processing (NLP), there remains a notable resource deficit for its diverse regional dialects, such as those spoken in Chittagong, Sylhet, and Barisal. These dialects, often considered unintelligible to speakers of Standard Bengali, pose challenges due to their unique grammatical structures and phonetic variations. Some linguists categorize them as distinct languages. To address this, we present ONUBAD, a large and freely available dataset for the automatic translation of Chittagong, Sylhet, and Barisal dialects into Standard Bangla using a Neural Machine Translation (NMT) system. ONUBAD provides a parallel corpus of 1540 words, 130 clauses, and 980 sentences per regional dialect and their standard counterparts along with English translation. The dataset includes metadata on phonetic variations and grammatical features, aiming to bridge the gap between standard and non-standard forms of Bangla. It serves as a valuable resource for researchers in NLP, dialect studies, and linguistic preservation, helping to develop more accurate and contextually relevant translation models. The dataset was collected between July and September 2024 from diverse sources such as books, websites, and regional people with the help of regional dialect specialists. It is hosted by the Department of Computer Science and Engineering, Jahangirnagar University, and is freely accessible at https://data.mendeley.com/datasets/6ft99kf89b/2 |
format | Article |
id | doaj-art-d1ffad4a0c6343a28dc8c705505d7fb6 |
institution | Kabale University |
issn | 2352-3409 |
language | English |
publishDate | 2025-02-01 |
publisher | Elsevier |
record_format | Article |
series | Data in Brief |
spelling | doaj-art-d1ffad4a0c6343a28dc8c705505d7fb62025-01-31T05:11:47ZengElsevierData in Brief2352-34092025-02-0158111276ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley DataNusrat Sultana0Rumana Yasmin1Bijon Mallik2Mohammad Shorif Uddin3Department of Computer Science and Engineering, Jahangirnagar University, Dhaka, Bangladesh; Department of Computer Science and Engineering, Bangladesh University of Business and Technology, Dhaka, Bangladesh; Corresponding author at: Department of Computer Science and Engineering, Jahangirnagar University, Dhaka, Bangladesh.Department of Computer Science and Engineering, Jahangirnagar University, Dhaka, Bangladesh; Department of Computer Science and Engineering, Bangladesh University of Professionals, Dhaka, BangladeshDepartment of Computer Science and Engineering, Bangladesh University of Engineering and Technology, Dhaka, Bangladesh; Department of Computer Science and Engineering, Bangladesh University of Business and Technology, Dhaka, BangladeshDepartment of Computer Science and Engineering, Jahangirnagar University, Dhaka, BangladeshDespite significant research on the Bangla language in Natural Language Processing (NLP), there remains a notable resource deficit for its diverse regional dialects, such as those spoken in Chittagong, Sylhet, and Barisal. These dialects, often considered unintelligible to speakers of Standard Bengali, pose challenges due to their unique grammatical structures and phonetic variations. Some linguists categorize them as distinct languages. To address this, we present ONUBAD, a large and freely available dataset for the automatic translation of Chittagong, Sylhet, and Barisal dialects into Standard Bangla using a Neural Machine Translation (NMT) system. ONUBAD provides a parallel corpus of 1540 words, 130 clauses, and 980 sentences per regional dialect and their standard counterparts along with English translation. The dataset includes metadata on phonetic variations and grammatical features, aiming to bridge the gap between standard and non-standard forms of Bangla. It serves as a valuable resource for researchers in NLP, dialect studies, and linguistic preservation, helping to develop more accurate and contextually relevant translation models. The dataset was collected between July and September 2024 from diverse sources such as books, websites, and regional people with the help of regional dialect specialists. It is hosted by the Department of Computer Science and Engineering, Jahangirnagar University, and is freely accessible at https://data.mendeley.com/datasets/6ft99kf89b/2http://www.sciencedirect.com/science/article/pii/S2352340925000083Neural machine translationNatural language processingRegional language datasetText translationStandard Bangla language |
spellingShingle | Nusrat Sultana Rumana Yasmin Bijon Mallik Mohammad Shorif Uddin ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data Data in Brief Neural machine translation Natural language processing Regional language dataset Text translation Standard Bangla language |
title | ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data |
title_full | ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data |
title_fullStr | ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data |
title_full_unstemmed | ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data |
title_short | ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data |
title_sort | onubad a comprehensive dataset for automated conversion of bangla regional dialects into standard bengali dialectmendeley data |
topic | Neural machine translation Natural language processing Regional language dataset Text translation Standard Bangla language |
url | http://www.sciencedirect.com/science/article/pii/S2352340925000083 |
work_keys_str_mv | AT nusratsultana onubadacomprehensivedatasetforautomatedconversionofbanglaregionaldialectsintostandardbengalidialectmendeleydata AT rumanayasmin onubadacomprehensivedatasetforautomatedconversionofbanglaregionaldialectsintostandardbengalidialectmendeleydata AT bijonmallik onubadacomprehensivedatasetforautomatedconversionofbanglaregionaldialectsintostandardbengalidialectmendeleydata AT mohammadshorifuddin onubadacomprehensivedatasetforautomatedconversionofbanglaregionaldialectsintostandardbengalidialectmendeleydata |