The regression analysis of group truncated data
Abstract
This thesis considers the regression modelling of grouped binary data that is subject
to truncation, and explores some general issues relating to truncation.
The likelihood for simple binary and ordinal models is developed and the statistical
behaviour of these models is explored. The models are found to be well
behaved. The efficiency of the truncated model is compared with that of conditional
logistic regression, a competing technique. It is found that the truncated
model is always more efficient but requires additional assumptions about the data
generation process to be applicable.
The estimation of the full sample size, N , before truncation occurs is considered,
in quite general regression models. The case where the covariate distribution is
discrete is first considered. This is extended to allow continuous covariates, and
the additional difficulties involved are explored. The issue of setting confidence
intervals for N is discussed. A simulation study is used to explore the methods
behaviour.
Next, the Bayesian analysis of truncated regression models is considered. The
use of the empirical distribution of the observed covariates to facilitate the analysis
is explored. The posterior distribution of the models parameters under this
approach is derived and a Gibbs sampling algorithm implemented to explore the
posterior. The convergence properties of the algorithm is considered, and the techniques
behaviour assessed in a small simulation study.
The effect of over-dispersion on the analysis of group truncated binary data
is considered. The available methods of introducing over-dispersion in clustered
binary data are discussed and it is argued that only random effects models provide
a viable approach. Parameter estimation in these models is derived via a marginal
likelihood. In addition a score test is constructed to test for the presence of random
effects in group truncated binary data. The methods performance is demonstrated
using a simulation study.
Finally, the use of the bootstrap to estimate the sampling distribution of parameter
estimates from truncated data is considered in an appendix. The inherent
limitations of using resampling methodologies to investigate truncated data is
demonstrated. It is shown that the nonparametric advantages of the bootstrap are
not realised with truncated data due to the lack of observations on the truncated
class.
Description
Keywords
Citation
Collections
Source
Type
Book Title
Entity type
Access Statement
License Rights
Restricted until
Downloads
File
Description