管道(软件)
计算机科学
三维建模
感知
人工智能
工程制图
集合(抽象数据类型)
领域(数学)
工程类
实体造型
三维模型
计算机视觉
数据建模
系统建模
特征(语言学)
作者
Mahammadismail Y Quadri,Shubham Karangale,Sudarshan Honnappanavar,Kushal Itnal,Niranjan Muchandi,Salma Shahpur
标识
DOI:10.1109/i3ctcon68242.2026.11507332
摘要
Single-image 3D modeling critically depends on the quality of 2D cues extracted from the input image. However, many existing reconstruction pipelines implicitly assume clean object silhouettes and fully visible instances, assumptions that rarely hold in real-world scenarios. To address this limitation, we propose a multi-task 2D perception suite designed to operate as a dedicated preprocessing layer for single-image 3D reconstruction systems. The proposed framework consists of four independently trained deep networks: a ResNet50-based classifier for predicting nine object categories from the Pix3D dataset, a DeepLabV3 model with a ResNet101 backbone for binary silhouette segmentation, and two additional ResNet50 networks for occlusion and truncation detectionAll models are trained using a unified strategy that incorporates strong regularization techniques and curriculum-style staged layer unfreezing to improve generalization. The independently optimized networks are integrated into a single inference pipeline that produces structured outputs including object class labels, pixel-accurate segmentation masks, occlusion and truncation probabilities, geometric border-touch indicators, and visualization overlays. The final output is a compact CSV/JSON representation that can be directly consumed by Pix3D-based 3D reconstruction methods, providing a reliable and interpretable 2D perception foundation for downstream 3D modeling tasks.
科研通智能强力驱动
Strongly Powered by AbleSci AI