计算机科学
人工智能
外部数据表示
一致性(知识库)
不变(物理)
代表(政治)
标记数据
编码器
数据一致性
机器学习
模式识别(心理学)
数据挖掘
数学
政治学
操作系统
法学
政治
数学物理
作者
Jun-Peng Fang,Caizhi Tang,Qing Cui,Feng Zhu,Longfei Li,Jun Zhou,Wei Zhu
标识
DOI:10.1145/3511808.3557699
摘要
Data augmentation-based semi-supervised learning (SSL) methods have made great progress in computer vision and natural language processing areas. One of the most important factors is that the semantic structure invariance of these data allows the augmentation procedure (e.g., rotating images or masking words) to thoroughly utilize the enormous amount of unlabeled data. However, the tabular data does not possess an obvious invariant structure, and therefore similar data augmentation methods do not apply to it. To fill this gap, we present a simple yet efficient data augmentation method particular designed for tabular data and apply it to the SSL algorithm: SDAT (Semi-supervised learning with Data Augmentation for Tabular data). We adopt a multi-task learning framework that consists of two components: the data augmentation procedure and the consistency training procedure. The data augmentation procedure which perturbs in latent space employs a variational auto-encoder (VAE) to generate the reconstructed samples as augmented samples. The consistency training procedure constrains the predictions to be invariant between the augmented samples and the corresponding original samples. By sharing a representation network (encoder), we jointly train the two components to improve effectiveness and efficiency. Extensive experimental studies validate the effectiveness of the proposed method on the tabular datasets.
科研通智能强力驱动
Strongly Powered by AbleSci AI