BanglaLM: Data Mining based Bangla Corpus for Language Model Research

作者
Md. Kowsher,Mohammed Jashim Uddin,Anik Tahabilder,Md. Ruhul Amin,Md. Fahim Shahriar,Md. Shohanur Islam Sobuj
出处
期刊:2021 Third International Conference on Inventive Research in Computing Applications (ICIRCA) 卷期号:5: 1435-1438 被引量:6
标识
DOI:10.1109/icirca51532.2021.9544818
摘要

Natural language processing (NLP) is an area of machine learning that has garnered a lot of attention in recent days due to the revolution in artificial intelligence, robotics, and smart devices. NLP focuses on training machines to understand and analyze various languages, extract meaningful information from those, translate from one language to another, correct grammar, predict the next word, complete a sentence, or even generate a completely new sentence from an existing corpus. A major challenge in NLP lies in training the model for obtaining high prediction accuracy since training needs a vast dataset. For widely used languages like English, there are many datasets available that can be used for NLP tasks like training a model and summarization but for languages like Bengali, which is only spoken primarily in South Asia, there is a dearth of big datasets which can be used to build a robust machine learning model. Therefore, NLP researchers who mainly work with the Bengali language will find an extensive, robust dataset incredibly useful for their NLP tasks involving the Bengali language. With this pressing issue in mind, this research work has prepared a dataset whose content is curated from social media, blogs, newspapers, wiki pages, and other similar resources. The amount of samples in this dataset is 19132010, and the length varies from 3 to 512 words. This dataset can easily be used to build any unsupervised machine learning model with an aim to performing necessary NLP tasks involving the Bengali language. Also, this research work is releasing two preprocessed version of this dataset that is especially suited for training both core machine learning-based and statistical-based model. As very few attempts have been made in this domain, keeping Bengali language researchers in mind, it is believed that the proposed dataset will significantly contribute to the Bengali machine learning and NLP community.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
骑猪兜风完成签到 ,获得积分10
2秒前
asdmwhx完成签到,获得积分10
3秒前
zzzy完成签到 ,获得积分10
3秒前
曼波曼波应助豆沙包采纳,获得10
5秒前
7秒前
11秒前
11秒前
11秒前
怪杰完成签到,获得积分10
13秒前
13秒前
lyk121438062完成签到,获得积分10
15秒前
科研助理795应助Taxwitted采纳,获得10
15秒前
zsh发布了新的文献求助10
16秒前
21秒前
飞快的蛋完成签到,获得积分0
22秒前
李鑫宁完成签到 ,获得积分10
22秒前
24秒前
Double_N完成签到,获得积分10
26秒前
爱上学的小金完成签到 ,获得积分10
27秒前
27秒前
小汪快跑完成签到 ,获得积分10
27秒前
潇洒的惋清应助王禄鑫采纳,获得10
28秒前
30秒前
gaowei完成签到 ,获得积分10
30秒前
31秒前
phdliuaccepted完成签到,获得积分10
32秒前
Copyright应助大头欢欢采纳,获得10
33秒前
不懂科研完成签到,获得积分10
33秒前
xishanju完成签到,获得积分10
34秒前
可爱的函函应助舒适香露采纳,获得10
34秒前
waswas完成签到,获得积分10
34秒前
38秒前
38秒前
陌桑子完成签到 ,获得积分10
39秒前
39秒前
温暖的寄容完成签到,获得积分10
39秒前
SilentLight完成签到,获得积分10
39秒前
火鸡味锅巴完成签到 ,获得积分10
40秒前
科研通AI6.3应助狂野悟空采纳,获得10
42秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Nondestructive Testing Handbook: Vol. 4, Thermal and Infrared Testing (IR), 4th ed 800
作者名:Kristopher P. Plain,悉尼大学的,目前只能查到其四篇论文,想找到其博士论文 590
Évora na Idade Média 555
Soil mites of the family Rhagidiidae (Actinedida: Eupodoidea). Morphology, Systematics, Ecology 520
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Radical Reactions 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7363851
求助须知:如何正确求助?哪些是违规求助? 8972897
关于积分的说明 19072450
捐赠科研通 7008781
什么是DOI,文献DOI怎么找? 3223773
关于科研通互助平台的介绍 2387472
邀请新用户注册赠送积分活动 2204605