Maximizing the Utility of Electronic Health Record Data for the Development, Evaluation, and Implementation of Real-time Prognostic Models
Electronic health record (EHR) data provide a rich resource for developing, evaluating, and implementing prognostic models to support real-time clinical decision-making. However, the complexity and incompleteness of EHR data present methodological and practical challenges that may limit prognostic performance. This dissertation develops approaches to maximize the utility of EHR data by addressing these challenges across the development, evaluation, and implementation of prognostic models.
The first chapter addresses model development by comparing conventional logistic regression with machine learning methods for prediction of venous thromboembolism among hospitalized adults using detailed EHR data on many risk factors. Although advances in machine learning and artificial intelligence have enabled increasingly complex prediction models, their performance relative to conventional statistical approaches remains unclear. This chapter compares the prediction performance of these methods to inform the appropriate modeling strategies for clinical risk prediction. The findings suggest that logistic regression performed as well as, or better than, the machine learning methods evaluated.
The second chapter focuses on model evaluation by developing a nonparametric estimator for the time-dependent area under the receiver operating characteristic curve (AUC) for continuous marks that quantify the severity of an event. Unlike the binary outcome models presented in the other chapters, this work considers time-to-event outcomes with continuous marks. A two-dimensional estimator for time and continuous marks was developed using the Nadaraya–Watson estimator with Gaussian kernels. Its asymptotic normality was established to construct confidence intervals for the AUC estimator. The proposed method is applied to hospital readmission data, with total cost per inpatient day serving as the continuous mark.
The third chapter addresses model implementation by evaluating strategies for handling informative presence, a common challenge in EHR data in which the collection of predictors is associated with patients' evolving clinical status. Through a simulation study, this chapter compares imputation strategies under varying degrees of informative presence and identifies combinations of imputation methods and prediction models that provide the most accurate and robust predictions. The findings suggest that multiple imputation by chained equations, combined with flexible machine learning algorithms, provides reliable prediction performance across a range of scenarios.
Collectively, these studies address approaches to leveraging EHR data on predictors and outcomes for developing and evaluating real-time prognostic models, while addressing the challenges in their implementation. These contributions strengthen the reliability and applicability of clinical prediction models and provide guidance for future research.