BACKGROUND: Osteosarcoma treatment decisions require accurate prognostic assessment, yet utilization of machine learning (ML) models is problematic because performance degrades consistently when models are applied across data sets. Current single-data-set models (single models) learn population-specific patterns rather than generalizable disease characteristics, limiting their clinical implementation. We developed a multicomponent-model (multi-model) ML framework using domain-adversarial training across 2 national registries, integrating structured clinical variables with text-based patient data to learn generalizable disease patterns and achieve reliable survival prediction across diverse populations.
METHODS: We conducted a retrospective study using data from 2 national cancer registries: SEER (Surveillance, Epidemiology, and End Results; n = 4,278 patients, 2004 to 2015) and NCDB (National Cancer Database; n = 4,049 patients, 2004 to 2018). We compared the cross-data-set performance of single models versus the multi-models. Primary outcomes were performance metrics for 2-year and 5-year overall survival predictions, measured by the area under the receiver operating characteristic curve (AUC), precision, recall, F1-score, and Brier score.
RESULTS: Single models achieved strong internal validation performance (AUC, 0.898 to 0.927) but performance declined substantially in cross-data-set validation (AUC, 0.563 to 0.665). The multi-model approach achieved cross-data-set AUCs of 0.708 to 0.843 for 2-year survival and 0.648 to 0.798 for 5-year survival, with improvements of 0.085 to 0.199 over single models across all evaluation metrics.
CONCLUSIONS: The multi-model ML approach demonstrated improved osteosarcoma prognostic ability across health-care data sets, addressing the generalizability challenges of models based on a single data set. Enhanced cross-data-set performance suggests the potential for consistent risk stratification to guide surgical planning, adjuvant therapy selection, and patient counseling. Prospective validation is needed to evaluate clinical impact.
LEVEL OF EVIDENCE: Prognostic Level III. See Instructions for Authors for a complete description of levels of evidence.
Read the article: PubMed · Publisher (DOI)