Summary Background The National Early Warning Score (NEWS) is widely used to predict patient deterioration, but retrospective validations face intervention bias and emphasise discrimination over clinical utility. We aimed to compare NEWS’s clinical net benefit against simplified scoring rules and machine learning, to determine whether simpler approaches suffice or if complex models offer meaningful advantages, and to evaluate performance heterogeneity across patient subgroups. Methods We included fifteen Danish hospitals with over 2·08 million hospital encounters representing 825200 unique patients over five years (2018 to 2023). We compared NEWS against both simpler and more complex approaches for predicting 24h mortality: Simplified NEWS (NEWS without blood pressure and temperature), DEWS (Simplified NEWS with age and sex), and a model based on eXtreme Gradient Boosting (XGB-EWS) incorporating vital signs, demographics, laboratory markers, plus medical history embeddings extracted using sentence transformers. We used propensity score weighting to mitigate intervention bias and evaluated performance using net benefit, Area Under the Receiver Operating Characteristic Curve (AUC), and calibration. Findings Decision curve analysis in the overall population showed maximum net benefit differences of 1·9 additional correct mortality identifications per 10000 patients between XGB-EWS and NEWS, and 1·5 per 10000 between NEWS and Simplified NEWS. However, stratified analyses revealed substantial heterogeneity: in the high-risk group (initial NEWS ≥7), XGB-EWS provided a maximum gain of 78·1 correct identifications per 10000 patients compared to NEWS, while Simplified NEWS resulted in a maximum performance loss of 60·2 per 10000. Conversely, in low-risk groups, differences were minimal. Interpretation Model complexity produced limited average clinical utility gains, concentrated among patients initially classified as high risk by NEWS. These findings support prospective trials of risk-stratified monitoring based on the first NEWS assessment, testing simplified approaches for lower-risk patients while retaining full NEWS or electronic health-record models for higher-risk patients. Funding Novo Nordisk Foundation Research in context Evidence before this study Early warning score systems have evolved from single vital sign monitoring to standardized multivariable scores such as the National Early Warning Score (NEWS), to even more sophisticated machine learning and deep learning frameworks utilizing electronic health records data. All these early warning scores are primarily designed to guide clinical decision-making by helping identify patients at risk of clinical deterioration within hospitals. Despite these advances, validation studies predominantly focus on statistical metrics measuring discrimination performance rather than meaningful clinical utility. On May 23, 2026 we searched PubMed for studies without language limitations published from May 23, 2016, to May 23, 2026, using terms including “early warning score”, “NEWS”, “NEWS2”, “machine learning”, “artificial intelligence”, “decision curve analysis”, “net benefit”, “clinical utility”, and “24-hour mortality.” While most of the published models, including the more sophisticated machine learning ones, have demonstrated better discrimination compared to traditional early warning scores, we found only one study that combined early warning score validation with clinical utility analysis for short-term clinical deterioration. Additionally, no studies have evaluated clinical utility across multiple patient subgroups while using causal inference approaches to address intervention bias in a large-scale healthcare system validation. Added value of this study The current work is a large-scale multi-center early warning score validation study, encompassing 2·08 million hospital encounters across 15 Danish hospitals and nine distinct clinical specialties over five years. We used the predictimand framework with causal inference methods to address intervention bias, a critical methodological advance in early warning score assessment. Unlike previous research focused primarily on discrimination metrics, we evaluated clinical utility using decision curve analysis across multiple patient subgroups, revealing heterogeneity in net benefit differences. In the overall population, sophisticated machine learning approaches provided marginal improvements (up to 1·9 per 10000 patients) despite superior discrimination, and simplified NEWS (without blood pressure and temperature) achieved comparable net benefit to full NEWS. However, in high-risk subgroups, particularly patients with high initial NEWS scores, differences increased substantially (up to 78·1 per 10000 for machine learning versus NEWS and 60·2 per 10000 for NEWS versus simplified NEWS), indicating that implementation strategies should be tailored to clinical context rather than uniformly applied. Implications of all the available evidence We demonstrated that differences in mortality prediction varied substantially by baseline risk: less than 2 per 10000 patients in the overall population, but up to 78 per 10000 in high-risk subgroups such as patients with initial high scores of NEWS. These findings challenge uniform implementation approaches: for general ward screening where most patients are at low baseline risk, simplified scoring achieves comparable utility while potentially enabling workflow optimization. However, for high-risk populations where baseline mortality is elevated, more complex models provide larger incremental gains that may justify additional implementation complexity. Future research should conduct prospective cluster-randomized trials evaluating context-specific implementation strategies, comparing simplified versus complex approaches across different care settings and baseline risk profiles.