Real-World Evidence and Machine Learning for Predictive Oncology and Clinical Decision Support: A Critical Narrative Review
Emmanuel Niiboye Odai *
Northeastern University, Massachusetts, United States.
Estherla Twene
Saint Luke’s Hospital, Kansas City, Missouri, United States.
Nurudeen Gbadegesin
University of Kentucky, Lexington, Kentucky, United States.
Andrew Oluwashijibomi Adegoju
Western Kentucky University, Kentucky, United States.
David Tetteh Akuaku Blemano
Michigan Technological University, Houghton, Michigan, United States.
Sampson Boateng
The Roux Institute - Northeastern University, Village Health Works, Maine, United States.
*Author to whom correspondence should be addressed.
Abstract
Real-world data from electronic health records, registries, claims, molecular testing and routine imaging have become central to contemporary oncology because they describe populations, treatments and outcomes that are incompletely represented in conventional trials. Machine learning can extend the utility of these data by extracting phenotypes from unstructured records, modelling prognosis, integrating clinicogenomic information and prioritising patients for clinical action. Yet the same combination creates a compound validity problem: routine-care data are generated by clinical processes rather than experimental design, while machine-learning models can amplify measurement error, confounding, selection effects and distribution shift. This critical narrative review evaluates the evidence linking real-world evidence and machine learning to predictive oncology and clinical decision support. Literature published from 1 January 2015 to 30 June 2026 was examined, with emphasis on peer-reviewed oncology studies, methodological guidance and prospective evaluations. The evidence is strongest for scalable extraction of treatment response, progression, performance status and mortality-related phenotypes, and for prognostic risk stratification using structured and narrative electronic health-record data. Evidence that prediction improves care is more limited. Randomised oncology studies show that machine-learning-triggered behavioural interventions can increase serious-illness conversations and reduce some forms of end-of-life treatment, but prospective evidence for treatment recommendation systems and large language model-enabled support remains dominated by concordance, simulation and workflow outcomes rather than patient benefit. Across the literature, external validation, calibration, outcome validity, transportability and separation of prognostic from causal questions are recurrent weaknesses. Real-world evidence can therefore be a powerful substrate and evaluation environment for predictive oncology, but data scale cannot substitute for fit-for-purpose measurement or causal design. Clinical translation should proceed through transparent reporting, independent validation, prospective workflow evaluation, equity assessment and continuous post-deployment monitoring, with human oversight retained for decisions in which model errors carry substantial clinical consequences.
Keywords: Real-world data, real-world evidence, machine learning, predictive oncology, clinical decision support, electronic health records, large language models, precision oncology