Improving Dietary Data Quality in Nutrition Studies: Development and Validation of an Explainable, Reproducible, Open Machine Learning-Assisted Framework for Outlier Detection in 24-Hour Recalls.
AI interpretation is pending for this paper.
Open original publication →What the AI sees
Not AI summarized yet.
Research significance
Pending deeper interpretation.
Source abstract
Twenty-four-hour dietary recalls (24HRs) are widely used in nutritional epidemiology, but repeated administration is required to capture within-person variability, introducing temporal dependence and challenges for data cleaning. Existing outlier detection approaches are largely static, manual, and based on adult-centric thresholds, limiting reproducibility and scalability. We developed and validated an explainable, reproducible machine learning (ML)-assisted framework to detect and evaluate implausible intake values in cross-sectional and longitudinal 24HRs and to generate AI-ready dietary datasets. The open pipeline integrates automated statistical detection using digitized National Cancer Institute 5th-95th percentile intake ranges, longitudinal interquartile range-based flagging, and ML-based rule explainability for outlier detection. Validation was conducted using the Microbiota, Growth, and Diet (MiGrowD) study in Ontario, Canada (126 children aged 8-12 years), and data from the 2015 Canadian Community Health Survey-Nutrition, a nationally representative sample of Canadian children and adults with repeated 24HRs (n = 15 216). Across energy intake and seven nutrients, 75 recalls containing at least one potentially implausible value were identified. In longitudinal data, the framework achieved > 77% sensitivity, 97% specificity, and 88% precision, outperforming the static method. This explainable data-driven framework improves data quality, transparency, and reproducibility, supporting scalable epidemiologic analyses using AI-ready dietary data.