计算机科学
条形码
稳健性(进化)
钥匙(锁)
编码器
光学字符识别
人工智能
推论
领域(数学分析)
文本识别
字符识别
计算机视觉
性格(数学)
智能字符识别
情报检索
模式识别(心理学)
可视化
数据挖掘
面部识别系统
搜索引擎索引
特征提取
面子(社会学概念)
文本检测
自然语言处理
作者
Chuanjun Chen,Yiru Yin,Junjie Liu
标识
DOI:10.1109/icairc68035.2025.11385200
摘要
Accurate cargo information acquisition is crucial for ensuring inventory precision and operational efficiency in automated storage and retrieval systems (AS/RS). Traditional methods including manual input, barcode scanning, and conventional optical character recognition (OCR) face significant limitations in complex warehouse environments with varying lighting conditions, text distortions, and domain-specific terminology. This paper proposes KAT-OCR (KnowledgeAugmented TrOCR), a novel text recognition framework that integrates domain knowledge with transformer-based OCR to address the challenges of pharmaceutical package text recognition. Our approach consists of two key innovations: the KnowledgeEnhanced Vision Encoder (KEVE) that dynamically fuses visual features with retrieved domain knowledge, and the KnowledgeConstrained Text Decoder (KCTD) that leverages domain constraints during text generation. Experimental results on the GY-PackInfo dataset demonstrate that KAT-OCR achieves 96.3% character accuracy and 81.7% key information accuracy, representing improvements of 3.2% and 9.2% respectively over vanilla TrOCR, while maintaining real-time inference speed (92 ms per image). The proposed method shows particular robustness in challenging scenarios including low-light conditions (+9.3% improvement), partial occlusion (+16.5% improvement), and confusing character recognition (+ 1 8. 4% improvement).
科研通智能强力驱动
Strongly Powered by AbleSci AI