Comparison of biostatistical and machine learning approaches to obtain prediction models with spike at zero variables as predictors

KeyNAKO-1142

Project leadProf. Dr. Heiko Becher

Approval date09.09.2025

Published date10.07.2026

SummaryThis methodological research project aims to compare different biostatistical and machine learning ap-proaches for developing prediction models in the presence of semi-continuous covariates (spike-at-zero variables), while accounting for additional confounders. The objective is to develop recommendations for prediction procedures in the presence of spike-at-zero variables. The analysis is restricted to a continuous outcome; for illustrative purposes, we use cholesterol level as the outcome variable. Spike-at-zero covari-ates include smoking (measured in pack-years) and alcohol consumption (measured in gram per day). Ad-ditional continuous covariates considered are age, BMI, body fat, and PHQ-9. Categorical or binary covari-ates include sex, diabetes status, socioeconomic status (SES), migrant status, and study center. The following statistical methods will be applied and compared. Full model, backward elimination, multivariable fractional polynomials (MFP), MFP with spike handling (MFPspike). Random forest is a popular machine learning approach and will be used here. An extension to binary outcomes or survival analysis is conceivable in a follow-up project.

Keywords Fractional-polynomials linear-regression machine-learning prediction-models random-forest spike-at-zero

InstitutionsUniversitätsklinikum Heidelberg, Medical Center-University of Freiburg Institute of Medical Biometry and Statistics: Universitatsklinikum Freiburg Institut fur Medizinische Biometrie und Statistik, Charité-Universitätsmedizin Berlin, Institute of Medical Biometry and Statistics, Medical Center - University of Freiburg

Go back