计算机科学
概括性
解析
人工智能
钥匙(锁)
表(数据库)
语言模型
统一模型
自然语言处理
工作流程
分离(微生物学)
编码(集合论)
情态动词
特征(语言学)
数据建模
源代码
编码(内存)
语义学(计算机科学)
机器学习
自然语言
作者
Wenwen Yu,Zhibo Yang,Jianqiang Wan,Sibo Song,Jun Tang,Wenqing Cheng,Yuliang Liu,Xiang Bai
标识
DOI:10.1109/tpami.2026.3677075
摘要
Visually-situated text parsing (VsTP) has recently seen notable advancements, driven by the growing demand for automated document understanding and the emergence of large language models capable of processing document-based questions. While various methods have been proposed to tackle the complexities of VsTP, existing solutions often rely on task-specific architectures and objectives for individual tasks. This leads to modal isolation and complex workflows due to the diversified targets and heterogeneous schemas. In this paper, we introduce OmniParser V2, a universal model that unifies VsTP typical tasks, including text spotting, key information extraction, table recognition, and layout analysis, into a unified framework. Central to our approach is the proposed Structured-Points-of-Thought (SPOT) prompting schemas, which improves model performance across diverse scenarios by leveraging a unified encoder-decoder architecture, objective, and input&output representation. SPOT eliminates the need for task-specific architectures and loss functions, significantly simplifying the processing pipeline. Our extensive evaluations across four tasks on eight different datasets show that OmniParser V2 achieves state-of-the-art or competitive results in VsTP. Additionally, we explore the integration of SPOT within a multimodal large language model structure, further enhancing visual text parsing capabilities on four tasks, thereby confirming the generality of SPOT prompting technique. The code is available at https://github.com/AlibabaResearch/AdvancedLiterateMachinery.
科研通智能强力驱动
Strongly Powered by AbleSci AI