计算机科学
自然语言处理
人工智能
语言识别
孟加拉语
鉴定(生物学)
脚本语言
语音识别
词(群论)
草书
自然语言
程序设计语言
语言学
哲学
生物
植物
作者
Muhammad Yasir,Li Chen,Amna Khatoon,Muhammad Amir Malik,Fazeel Abid
摘要
Mixed script identification is a hindrance for automated natural language processing systems. Mixing cursive scripts of different languages is a challenge because NLP methods like POS tagging and word sense disambiguation suffer from noisy text. This study tackles the challenge of mixed script identification for mixed‐code dataset consisting of Roman Urdu, Hindi, Saraiki, Bengali, and English. The language identification model is trained using word vectorization and RNN variants. Moreover, through experimental investigation, different architectures are optimized for the task associated with Long Short‐Term Memory (LSTM), Bidirectional LSTM, Gated Recurrent Unit (GRU), and Bidirectional Gated Recurrent Unit (Bi‐GRU). Experimentation achieved the highest accuracy of 90.17 for Bi‐GRU, applying learned word class features along with embedding with GloVe. Moreover, this study addresses the issues related to multilingual environments, such as Roman words merged with English characters, generative spellings, and phonetic typing.
科研通智能强力驱动
Strongly Powered by AbleSci AI