Machine Learning Modeling for Predicting Mortality in Pediatric Patients Undergoing Elective Noncardiac Surgery: Comparison to a Regression Model.
AI interpretation is pending for this paper.
Open original publication →What the AI sees
Not AI summarized yet.
Research significance
Pending deeper interpretation.
Source abstract
BACKGROUND: Perioperative mortality in children is relatively rare; however, accurate preoperative risk stratification is critical, as it enables anticipatory planning to mitigate the risk of death. This study aims to use machine learning (ML) to develop and internally validate a predictive model for 30-day mortality in children undergoing noncardiac surgery and compare model performance to the regression-based Pediatric Risk Assessment (PRAm) score. METHODS: A retrospective study of the National Surgical Quality Improvement Program (NSQIP)-Pediatric database from 2012 to 2022, excluding 2020, was performed. Patients <18 years undergoing multispecialty surgical procedures except cardiac surgery were included. Clinically meaningful risk factors for mortality were included in the random forest and XGBoost ML models. The primary outcome was 30-day mortality. RESULTS: A total of 1,023,639 unique patient encounters were included in the final analysis; 3522 (0.34%) resulted in 30-day mortality. Most patients of the 1,023,639 were ≥12 years (307,930, 30.1%), followed by 6 to 12 years (263,241, 25.7%). The majority was inpatient (596,642, 58.3%) and underwent an elective procedure (736,163, 71.9%). The most common comorbid conditions were neurologic disease (209,384, 20.5%), gastrointestinal disease (177,888, 17.4%), and central nervous system tumor or acquired abnormality (129,711, 12.7%). A total of 3.2% (33,103) were mechanically ventilated and 0.6% (6032) supported with inotropes. ML models were developed on the 70% training set (n = 716,662) and evaluated using the 30% validation set (n = 306,977). The XGBoost model demonstrated the best performance in the validation set (area under the receiver operating characteristic curve [AUC-ROC] = 0.956, area under the precision-recall curve [AUC-PR] = 0.179). The accuracy was 99.4% and the precision was 0.247, meaning that a positive prediction was associated with a 24.7% risk of mortality. The model demonstrated good calibration (Brier score = 0.003) between observed and expected probabilities. The AUC-ROC for the XGBoost model was 0.956 and for the PRAm score 0.958. There was no substantial increase in net benefit of the XGBoost model versus the PRAm score across the range of threshold probabilities. CONCLUSIONS: ML can be leveraged to develop clinical prediction tools with excellent predictive performance for rare but critical outcomes like postoperative mortality. However, the ML-based models in the current study performed similarly to the regression-based PRAm score, highlighting the need to consider whether their added complexity yields meaningful clinical benefit.