Improving Dietary Data Quality in Nutrition Studies: Development and Validation of an Explainable, Reproducible, Open Machine Learning-Assisted Framework for Outlier Detection in 24-Hour Recalls.
The paper reports development and validation of an open, explainable machine-learning-assisted framework that identified potentially implausible values in repeated 24-hour dietary recalls and improved outlier-detection performance over a static method in Canadian child and adult datasets.
Open original publication →What the AI sees
The paper reports development and validation of an open, explainable machine-learning-assisted framework that identified potentially implausible values in repeated 24-hour dietary recalls and improved outlier-detection performance over a static method in Canadian child and adult datasets.
Research significance
Evidence: the framework improves cleaning of longitudinal dietary-recall data, with reported sensitivity above 77%, specificity of 97%, and precision of 88%; inference: it could indirectly strengthen nutrition research relevant to pediatric oncology, but this record provides no cancer-specific application, treatment hypothesis, or evidence of improved clinical outcomes.
Source abstract
Twenty-four-hour dietary recalls (24HRs) are widely used in nutritional epidemiology, but repeated administration is required to capture within-person variability, introducing temporal dependence and challenges for data cleaning. Existing outlier detection approaches are largely static, manual, and based on adult-centric thresholds, limiting reproducibility and scalability. We developed and validated an explainable, reproducible machine learning (ML)-assisted framework to detect and evaluate implausible intake values in cross-sectional and longitudinal 24HRs and to generate AI-ready dietary datasets. The open pipeline integrates automated statistical detection using digitized National Cancer Institute 5th-95th percentile intake ranges, longitudinal interquartile range-based flagging, and ML-based rule explainability for outlier detection. Validation was conducted using the Microbiota, Growth, and Diet (MiGrowD) study in Ontario, Canada (126 children aged 8-12 years), and data from the 2015 Canadian Community Health Survey-Nutrition, a nationally representative sample of Canadian children and adults with repeated 24HRs (n = 15 216). Across energy intake and seven nutrients, 75 recalls containing at least one potentially implausible value were identified. In longitudinal data, the framework achieved > 77% sensitivity, 97% specificity, and 88% precision, outperforming the static method. This explainable data-driven framework improves data quality, transparency, and reproducibility, supporting scalable epidemiologic analyses using AI-ready dietary data.