特征选择
计算机科学
特征(语言学)
维数之咒
错误发现率
数据挖掘
构造(python库)
统计的
集合(抽象数据类型)
选择(遗传算法)
任务(项目管理)
信息隐私
数据集
人工智能
概率分布
机器学习
特征提取
选型
功率(物理)
方案(数学)
合成数据
分布(数学)
统计模型
控制(管理)
变量(数学)
模式识别(心理学)
利用
顺序统计量
作者
Jie Hu,Jiayi Tong,Yang Ning,Cheng Yong Tang,Jason H. Moore,Runze Li,Yong Chen
标识
DOI:10.1093/jrsssb/qkaf074
摘要
Abstract Selecting a set of universally relevant features associated with a given response variable across multiple distributed data sites is an important problem in numerous scientific fields. However, performing this federated feature selection task becomes challenging when individual-level data cannot be shared due to privacy concerns. The problem is further complicated by potential heterogeneity in both feature distributions and model parameters across sites. In this paper, we propose Fed-false discovery rate (FDR), a federated feature selection framework that simultaneously identifies important features while controlling the FDR. To ensure privacy preservation and reduce communication costs, the Fed-FDR shares only lower-dimensional coefficient estimates instead of transmitting summary statistics for all features, with the dimensionality shown to be of the same order as the number of relevant features. The coordinating centre then leverages these lower-dimensional coefficient estimates to construct a generalized mirror statistic to identify the important features. The Fed-FDR is robust to the heterogeneity of feature distribution and model parameters, easy to implement, and computationally efficient. We further demonstrate that Fed-FDR effectively controls the FDR while achieving strong statistical power in our simulation studies. The results of the empirical study also demonstrate that the method is both valid and implementation-ready.
科研通智能强力驱动
Strongly Powered by AbleSci AI