计算机科学
光学字符识别
文件处理
文本识别
过程(计算)
特征提取
人工智能
特征(语言学)
图像处理
情报检索
分割
性格(数学)
智能字符识别
字符识别
图像(数学)
自然语言处理
语言学
几何学
哲学
操作系统
数学
作者
Ridvy Avyodri,Samuel Lukas,Hendra Tjahyadi
标识
DOI:10.1109/ictiia54654.2022.9935961
摘要
Most organizations worldwide still rely on paper-based documents. Usage of paper-based documents gives a hard time extracting the data required from those documents. This heavy paper usage also damages the efficiency in cost and time, not to mention the impact on the environment caused by deforestation to produce these papers. These are some reasons that motivate the need to digitalize paper-based documents. To convert the usage of paper-based documents into paperless documents cannot be done in an instant. In its transition, these paper-based documents are usually scanned into image format to reduce the usage of paper. From this comes a need for technology that is able to recognize and extract data in the scanned image of paper-based documents. Optical Character Recognition makes it possible to do text recognition appearing in images. However, despite its long history of development, OCR for text recognition has yet to achieve 100% accuracy. In general, OCR process will be divided into Image Pre-processing, Text Segmentation/Localization, Feature Extraction, Text Recognition, and Post-Processing. Thus, this research will review OCR-related works and the methods used within this framework to support further research.
科研通智能强力驱动
Strongly Powered by AbleSci AI