ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data

Despite significant research on the Bangla language in Natural Language Processing (NLP), there remains a notable resource deficit for its diverse regional dialects, such as those spoken in Chittagong, Sylhet, and Barisal. These dialects, often considered unintelligible to speakers of Standard Benga...

Full description

Saved in:

Bibliographic Details
Main Authors:	Nusrat Sultana, Rumana Yasmin, Bijon Mallik, Mohammad Shorif Uddin
Format:	Article
Language:	English
Published:	Elsevier 2025-02-01
Series:	Data in Brief
Subjects:	Neural machine translation Natural language processing Regional language dataset Text translation Standard Bangla language
Online Access:	http://www.sciencedirect.com/science/article/pii/S2352340925000083
Tags:	Add Tag No Tags, Be the first to tag this record!

_version_	1832576508927934464
author	Nusrat Sultana Rumana Yasmin Bijon Mallik Mohammad Shorif Uddin
author_facet	Nusrat Sultana Rumana Yasmin Bijon Mallik Mohammad Shorif Uddin
author_sort	Nusrat Sultana
collection	DOAJ
description	Despite significant research on the Bangla language in Natural Language Processing (NLP), there remains a notable resource deficit for its diverse regional dialects, such as those spoken in Chittagong, Sylhet, and Barisal. These dialects, often considered unintelligible to speakers of Standard Bengali, pose challenges due to their unique grammatical structures and phonetic variations. Some linguists categorize them as distinct languages. To address this, we present ONUBAD, a large and freely available dataset for the automatic translation of Chittagong, Sylhet, and Barisal dialects into Standard Bangla using a Neural Machine Translation (NMT) system. ONUBAD provides a parallel corpus of 1540 words, 130 clauses, and 980 sentences per regional dialect and their standard counterparts along with English translation. The dataset includes metadata on phonetic variations and grammatical features, aiming to bridge the gap between standard and non-standard forms of Bangla. It serves as a valuable resource for researchers in NLP, dialect studies, and linguistic preservation, helping to develop more accurate and contextually relevant translation models. The dataset was collected between July and September 2024 from diverse sources such as books, websites, and regional people with the help of regional dialect specialists. It is hosted by the Department of Computer Science and Engineering, Jahangirnagar University, and is freely accessible at https://data.mendeley.com/datasets/6ft99kf89b/2
format	Article
id	doaj-art-d1ffad4a0c6343a28dc8c705505d7fb6
institution	Kabale University
issn	2352-3409
language	English
publishDate	2025-02-01
publisher	Elsevier
record_format	Article
series	Data in Brief
spelling	doaj-art-d1ffad4a0c6343a28dc8c705505d7fb62025-01-31T05:11:47ZengElsevierData in Brief2352-34092025-02-0158111276ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley DataNusrat Sultana0Rumana Yasmin1Bijon Mallik2Mohammad Shorif Uddin3Department of Computer Science and Engineering, Jahangirnagar University, Dhaka, Bangladesh; Department of Computer Science and Engineering, Bangladesh University of Business and Technology, Dhaka, Bangladesh; Corresponding author at: Department of Computer Science and Engineering, Jahangirnagar University, Dhaka, Bangladesh.Department of Computer Science and Engineering, Jahangirnagar University, Dhaka, Bangladesh; Department of Computer Science and Engineering, Bangladesh University of Professionals, Dhaka, BangladeshDepartment of Computer Science and Engineering, Bangladesh University of Engineering and Technology, Dhaka, Bangladesh; Department of Computer Science and Engineering, Bangladesh University of Business and Technology, Dhaka, BangladeshDepartment of Computer Science and Engineering, Jahangirnagar University, Dhaka, BangladeshDespite significant research on the Bangla language in Natural Language Processing (NLP), there remains a notable resource deficit for its diverse regional dialects, such as those spoken in Chittagong, Sylhet, and Barisal. These dialects, often considered unintelligible to speakers of Standard Bengali, pose challenges due to their unique grammatical structures and phonetic variations. Some linguists categorize them as distinct languages. To address this, we present ONUBAD, a large and freely available dataset for the automatic translation of Chittagong, Sylhet, and Barisal dialects into Standard Bangla using a Neural Machine Translation (NMT) system. ONUBAD provides a parallel corpus of 1540 words, 130 clauses, and 980 sentences per regional dialect and their standard counterparts along with English translation. The dataset includes metadata on phonetic variations and grammatical features, aiming to bridge the gap between standard and non-standard forms of Bangla. It serves as a valuable resource for researchers in NLP, dialect studies, and linguistic preservation, helping to develop more accurate and contextually relevant translation models. The dataset was collected between July and September 2024 from diverse sources such as books, websites, and regional people with the help of regional dialect specialists. It is hosted by the Department of Computer Science and Engineering, Jahangirnagar University, and is freely accessible at https://data.mendeley.com/datasets/6ft99kf89b/2http://www.sciencedirect.com/science/article/pii/S2352340925000083Neural machine translationNatural language processingRegional language datasetText translationStandard Bangla language
spellingShingle	Nusrat Sultana Rumana Yasmin Bijon Mallik Mohammad Shorif Uddin ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data Data in Brief Neural machine translation Natural language processing Regional language dataset Text translation Standard Bangla language
title	ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data
title_full	ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data
title_fullStr	ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data
title_full_unstemmed	ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data
title_short	ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data
title_sort	onubad a comprehensive dataset for automated conversion of bangla regional dialects into standard bengali dialectmendeley data
topic	Neural machine translation Natural language processing Regional language dataset Text translation Standard Bangla language
url	http://www.sciencedirect.com/science/article/pii/S2352340925000083
work_keys_str_mv	AT nusratsultana onubadacomprehensivedatasetforautomatedconversionofbanglaregionaldialectsintostandardbengalidialectmendeleydata AT rumanayasmin onubadacomprehensivedatasetforautomatedconversionofbanglaregionaldialectsintostandardbengalidialectmendeleydata AT bijonmallik onubadacomprehensivedatasetforautomatedconversionofbanglaregionaldialectsintostandardbengalidialectmendeleydata AT mohammadshorifuddin onubadacomprehensivedatasetforautomatedconversionofbanglaregionaldialectsintostandardbengalidialectmendeleydata

ONUBAD: A comprehensive dataset for automated conversion of Bangla regional dialects into standard Bengali dialectMendeley Data

Similar Items