计算机科学
人工智能
背景(考古学)
语音识别
自然语言处理
模式识别(心理学)
机器学习
生物
古生物学
作者
Xiaomeng Yang,Zhi Qiang Qiao,Wei Jin,Dongbao Yang,Yu Zhou
标识
DOI:10.1109/lsp.2024.3381893
摘要
Scene Text Recognition (STR) is challenging because of various text styles, shapes, and backgrounds. Although the integration of linguistic information enhances models' performance, existing methods based on either permuted language modeling (PLM) or masked language modeling (MLM) have their drawbacks. PLM's autoregressive decoding lacks foresight into subsequent characters, while MLM overlooks intercharacter dependencies. To address these problems, we propose a masked and permuted implicit context learning network for STR, which unifies PLM and MLM within a single decoder, inheriting the advantages of both approaches. We utilize the training procedure of PLM and incorporate word length information into the decoding process to integrate MLM, substituting the undetermined characters with mask tokens. Besides, we employ the perturbation training technique to train a more robust model against potential length prediction errors. Our comprehensive evaluations demonstrate the performance of our model. It achieves superior performance on the popularly used benchmarks and outperforms previous state-of-the-art methods with a substantial improvement of 9.1% on the more challenging Union14M-Benchmark.
科研通智能强力驱动
Strongly Powered by AbleSci AI