已入深夜,您辛苦了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!祝你早点完成任务,早点休息,好梦!

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

计算机科学 概括性 解析 人工智能 钥匙(锁) 表(数据库) 语言模型 统一模型 自然语言处理 工作流程 分离(微生物学) 编码(集合论) 情态动词 特征(语言学) 数据建模 源代码 编码(内存) 语义学(计算机科学) 机器学习 自然语言
作者
Wenwen Yu,Zhibo Yang,Jianqiang Wan,Sibo Song,Jun Tang,Wenqing Cheng,Yuliang Liu,Xiang Bai
出处
期刊:IEEE Transactions on Pattern Analysis and Machine Intelligence [IEEE Computer Society]
卷期号:PP: 1-18 被引量:1
标识
DOI:10.1109/tpami.2026.3677075
摘要

Visually-situated text parsing (VsTP) has recently seen notable advancements, driven by the growing demand for automated document understanding and the emergence of large language models capable of processing document-based questions. While various methods have been proposed to tackle the complexities of VsTP, existing solutions often rely on task-specific architectures and objectives for individual tasks. This leads to modal isolation and complex workflows due to the diversified targets and heterogeneous schemas. In this paper, we introduce OmniParser V2, a universal model that unifies VsTP typical tasks, including text spotting, key information extraction, table recognition, and layout analysis, into a unified framework. Central to our approach is the proposed Structured-Points-of-Thought (SPOT) prompting schemas, which improves model performance across diverse scenarios by leveraging a unified encoder-decoder architecture, objective, and input&output representation. SPOT eliminates the need for task-specific architectures and loss functions, significantly simplifying the processing pipeline. Our extensive evaluations across four tasks on eight different datasets show that OmniParser V2 achieves state-of-the-art or competitive results in VsTP. Additionally, we explore the integration of SPOT within a multimodal large language model structure, further enhancing visual text parsing capabilities on four tasks, thereby confirming the generality of SPOT prompting technique. The code is available at https://github.com/AlibabaResearch/AdvancedLiterateMachinery.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
2秒前
奥奥脑袋发布了新的文献求助10
2秒前
眯眯眼的安雁完成签到 ,获得积分10
3秒前
4秒前
GUI完成签到,获得积分10
5秒前
四季大枣发布了新的文献求助10
5秒前
happy完成签到 ,获得积分10
6秒前
SCI完成签到,获得积分10
6秒前
时间太少了完成签到,获得积分10
6秒前
个性大白完成签到,获得积分20
7秒前
7秒前
7秒前
陈凯鸿发布了新的文献求助10
8秒前
雨田发布了新的文献求助10
11秒前
个性大白发布了新的文献求助10
11秒前
坚定青槐完成签到 ,获得积分10
11秒前
奥奥脑袋完成签到,获得积分10
13秒前
14秒前
hai完成签到,获得积分10
15秒前
zephyr的应助被SHAN采纳,获得10
15秒前
maner完成签到 ,获得积分10
16秒前
Akim的应助被943034197采纳,获得10
17秒前
17秒前
香蕉觅云的应助被张陶求采纳,获得10
18秒前
小二郎的应助被Rainandbow采纳,获得10
18秒前
19秒前
我是老大的应助被科研通管家采纳,获得10
19秒前
李健的应助被科研通管家采纳,获得10
20秒前
20秒前
香蕉觅云的应助被科研通管家采纳,获得10
20秒前
深情安青的应助被科研通管家采纳,获得10
20秒前
隐形曼青的应助被科研通管家采纳,获得10
20秒前
研友_VZG7GZ的应助被科研通管家采纳,获得10
20秒前
深情安青的应助被科研通管家采纳,获得10
20秒前
SciGPT的应助被科研通管家采纳,获得10
20秒前
20秒前
搜集达人的应助被科研通管家采纳,获得10
20秒前
21秒前
斯文败类的应助被科研通管家采纳,获得10
21秒前
不会科研的大锤完成签到,获得积分20
21秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Rosenblum, Global Change Biology 800
自動車の空力技術 800
Using Projective Methods with Children 600
Organizational Behavior 510
Management and the Arts 510
Issues in Task-Based Language Teaching 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7784985
求助须知:如何正确求助?哪些是违规求助? 9324126
关于积分的说明 20397323
捐赠科研通 7373621
什么是DOI,文献DOI怎么找? 3321217
关于科研通互助平台的介绍 2469095
邀请新用户注册赠送积分活动 2337487