缺少数据
插补(统计学)
范畴变量
扩展(谓词逻辑)
计算机科学
数学
数据挖掘
统计
程序设计语言
作者
Kristin N. Javaras,David A. van Dyk
出处
期刊:
日期:2003-09-01
卷期号:98 (463): 703-715
被引量:19
标识
DOI:10.1198/016214503000000611
摘要
AbstractWe consider the application of multiple imputation to data containing not only partially missing categorical and continuous variables, but also partially missing 'semicontinuous' variables (variables that take on a single discrete value with positive probability but are otherwise continuously distributed). As an imputation model for data sets of this type, we introduce an extension of the standard general location model proposed by Olkin and Tate; our extension, the blocked general location model, provides a robust and general strategy for handling partially observed semicontinuous variables. In particular, we incorporate a two-level model for the semicontinuous variables into the general location model. The first level models the probability that the semicontinuous variable takes on its point mass value, and the second level models the distribution of the variable given that it is not at its point mass. In addition, we introduce EM and data augmentation algorithms for the blocked general location model with missing data; these can be used to generate imputations under the proposed model and have been implemented in publicly available software. We illustrate our model and computational methods via a simulation study and an analysis of a survey of Massachusetts Megabucks Lottery winners.KEY WORDS: Data augmentationEM algorithmGeneral location modelMissing dataSurvey data
科研通智能强力驱动
Strongly Powered by AbleSci AI