Counting co-occurring diseases to predict mortality is as accurate as multimorbidity indices: an external validation study
Counting a person’s chronic diseases predicted death about as well as seven more complex scoring methods.
*Prospective cohort external validation study; Level 1b (OCEBM).
*Prospective cohort external validation study; Level 1b (OCEBM).
Citation
Velek P, Splinter MJ, Stirland L, et al. Counting co-occurring diseases to predict mortality is as accurate as multimorbidity indices: an external validation study. Age and Ageing. 2026;55:afag150. doi:10.1093/ageing/afag150.
Background
Many tools combine multiple chronic conditions into a single score to estimate a patient’s risk of death, but it is unclear whether they add value beyond simply counting conditions. This study directly compared seven recommended scoring methods against disease-count approaches in the same populations.
Patients
Community-dwelling adults from the Rotterdam Study (Netherlands), with seven index-specific sub-cohorts (n=2,409 to 9,045). Some sub-cohorts were restricted by age or sex based on the original tool; participants missing required baseline data were excluded.
Intervention
Seven published multimorbidity scoring methods used as predictors of all-cause death.
Control
(1) Age and sex only; (2) age, sex, and a simple count of 10 chronic diseases; plus other count-based benchmarks.
Outcome
Ability to predict all-cause death (discrimination and calibration).
Follow-up Period
1 to 6 years (varied by sub-cohort).
Results
| Finding across 7 sub-cohorts | What was observed |
|---|---|
| Discrimination (primary) | Simple disease counts and complex scores were nearly identical (maximum difference in C-statistic 0.06); both were better than age and sex alone. |
| Overall prediction accuracy | Similar across models; maximum improvement in Brier score about 4%. |
| Calibration (agreement of predicted vs observed risk) | Poor for 4 of 7 complex scores, especially those with follow-up of 2 years or less. |
C-statistic ranges from 0.5 (no separation) to 1.0 (perfect separation).
Limitations
Some original score components could not be perfectly recreated (for example, certain diagnoses and diabetes severity). Several tools were developed in different settings (including post-hospital discharge), which may explain miscalibration. Differences in prediction performance were small and may not meaningfully change individual patient decisions.
Funding
Dutch public funders and European support; funders had no role.
Clinical Application
For community-dwelling older adults, use a straightforward chronic disease count (with age/sex) for mortality risk discussions; complex scores add little and may misestimate risk.
Discussion
Sign in to join the discussion.
In this external validation study using the prospective Rotterdam Study, multimorbidity indices were no better than a simple count of co-occurring diseases for predicting all-cause mortality (maximum C-statistic difference 0.06). Given poor calibration in 4/7 indices, would you change practice toward disease counts, or does miscalibration/transportability limit clinical use? The authors compared models using the C-statistic (e.g., C-statistic ~0.80–0.82 in several samples) and noted that some indices were miscalibrated despite similar C-statistics. What does the C-statistic primarily measure in this context?