人工智能
模式识别(心理学)
特征(语言学)
计算机科学
启发式
特征向量
管道(软件)
像素
特征提取
歧管(流体力学)
上下文图像分类
钥匙(锁)
对象(语法)
非线性降维
忠诚
图像(数学)
视觉对象识别的认知神经科学
人工神经网络
歧管对齐
数据建模
最优化问题
机器学习
卷积神经网络
目标检测
计算机视觉
高斯分布
数据分类
欧几里德距离
降维
直方图
统计分类
欧几里得空间
数学
数据挖掘
特征检测(计算机视觉)
高斯过程
特征学习
作者
Fangqing Liu,Han Huang,Fujian Feng,Xueming Yan,Zhifeng Hao
标识
DOI:10.1109/tpami.2026.3657249
摘要
Data augmentation is crucial for addressing insufficient training data, especially for augmenting positive samples. However, existing methods mostly rely on neural network-based feedback for data augmentation and often overlook the optimization of feature distribution. In this study, we present a practical, distribution-preserving data augmentation pipeline that augments positive samples by optimizing a feature indicator (e.g., two-dimensional entropy), aiming to maintain alignment with the original data distribution. Inspired by the manifold hypothesis, we propose a Manifold Heuristic Optimization Algorithm (MHOA), which augments positive samples by exploring the low-dimensional Euclidean space around object contour pixels instead of the entire decision space. Guided by a "distribution-preservation-first" perspective, our approach explicitly optimizes fidelity to the original data manifold and only retains augmented samples whose feature statistics (e.g., mean, variance) align with the source class. It significantly improves image classification accuracy across neural networks, outperforming state-of-the-art data augmentation methods-especially when the dataset's feature indicator follows a Gaussian distribution. The algorithm's search space, focused on neighborhoods of key feature pixels, is the core driver of its superior performance.
科研通智能强力驱动
Strongly Powered by AbleSci AI