计算机科学
编码(内存)
人工智能
过程(计算)
编码器
人类视觉系统模型
视觉推理
图像(数学)
计算机视觉
认知
因果推理
因果模型
语义学(计算机科学)
视觉感受
可视化
计算模型
光流
认知科学
自然语言处理
人机交互
钥匙(锁)
工作(物理)
机制(生物学)
图像处理
模式识别(心理学)
动作(物理)
因果关系(物理学)
订单(交换)
作者
Haoran Wei,Yaofeng Sun,Yukun Li
标识
DOI:10.48550/arxiv.2601.20552
摘要
We present DeepSeek-OCR 2 to investigate the feasibility of a novel encoder-DeepEncoder V2-capable of dynamically reordering visual tokens upon image semantics. Conventional vision-language models (VLMs) invariably process visual tokens in a rigid raster-scan order (top-left to bottom-right) with fixed positional encoding when fed into LLMs. However, this contradicts human visual perception, which follows flexible yet semantically coherent scanning patterns driven by inherent logical structures. Particularly for images with complex layouts, human vision exhibits causally-informed sequential processing. Inspired by this cognitive mechanism, DeepEncoder V2 is designed to endow the encoder with causal reasoning capabilities, enabling it to intelligently reorder visual tokens prior to LLM-based content interpretation. This work explores a novel paradigm: whether 2D image understanding can be effectively achieved through two-cascaded 1D causal reasoning structures, thereby offering a new architectural approach with the potential to achieve genuine 2D reasoning. Codes and model weights are publicly accessible at http://github.com/deepseek-ai/DeepSeek-OCR-2.
科研通智能强力驱动
Strongly Powered by AbleSci AI