计算机科学
判别式
先验概率
人工智能
情态动词
机器学习
数据挖掘
模式识别(心理学)
贝叶斯概率
化学
高分子化学
作者
Yupeng Song,Naifu Liang,Qing Guo,Jicheng Dai,Junwei Bai,Fazhi He
标识
DOI:10.1016/j.ipm.2023.103497
摘要
Text, 2D, and 3D information are crucial information representations in modern science and management disciplines. However, complex and irregular 3D data produce data scarcity and expensive generation that limit their processing and application. In this paper, we present MeshCLIP, a new cross-modal information learning paradigm to directly process 3D mesh data end-to-end in a zero/few-shot manner. Specifically, we design a novel pipeline based on visual factors and graphics principles, bridging the gap between 3D mesh data and other modal data, thereby joining 2/3D visual and textual information for zero/few-shot learning. Then, we construct a self-attention adapter for 3D mesh key information learning when training only a few priors, significantly improving the model’s discriminative ability. Extensive experiments demonstrate that the proposed MeshCLIP can achieve state-of-the-art results on multiple challenging 3D mesh datasets. In the whole 3D domain, the proposed zero-shot approach significantly outperforms the existing other 3D representation methods with an accuracy 3 × better ( increased by 41.5%) on the ModelNet40 dataset. Furthermore, in few-shot learning, the proposed MeshCLIP uses only a few supervised priors (only less than 10% of the sample size) to achieve results close to those of methods trained on a full dataset.
科研通智能强力驱动
Strongly Powered by AbleSci AI