计算机科学
感知
解析
理解力
领域(数学分析)
人工智能
自然语言处理
空格(标点符号)
对偶(语法数字)
人机交互
机器学习
语言学
数学分析
程序设计语言
神经科学
哲学
操作系统
生物
数学
作者
Hao Feng,Qi Liu,Hao Liu,Wengang Zhou,Houqiang Li,Can Huang
标识
DOI:10.48550/arxiv.2311.11810
摘要
This work presents DocPedia, a novel large multimodal model (LMM) for versatile OCR-free document understanding, capable of parsing images up to 2,560$\times$2,560 resolution. Unlike existing work either struggle with high-resolution documents or give up the large language model thus vision or language ability constrained, our DocPedia directly processes visual input in the frequency domain rather than the pixel space. The unique characteristic enables DocPedia to capture a greater amount of visual and textual information using a limited number of visual tokens. To consistently enhance both perception and comprehension abilities of our model, we develop a dual-stage training strategy and enrich instructions/annotations of all training tasks covering multiple document types. Extensive quantitative and qualitative experiments conducted on various publicly available benchmarks confirm the mutual benefits of jointly learning perception and comprehension tasks. The results provide further evidence of the effectiveness and superior performance of our DocPedia over other methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI