Scaling Early Literacy Screening for Sustainable Education: A Cloud-Native Architecture Integrating Machine Learning and Human-in-the-Loop Validation
Early literacy screening is essential for reducing long-term educational inequality, yet traditional paper-based assessments remain difficult to scale due to logistical constraints and delayed feedback. This study presents K-KOBUKI, a cloud-based prototype screening workflow that organizes early literacy assessment as a human-validated, data-driven process. The system integrates structured assessment responses with automated speech recognition-based analysis of oral reading performance across five literacy domains and incorporates a human-in-the-loop verification stage to ensure the reliability of speech-derived features. The system was evaluated using data from 195 first-grade students. Across repeated stratified cross-validation, multiple classification models achieved stable recall (≈0.85) under class imbalance conditions, supporting consistent identification of at-risk learners. Psychometric-informed feature refinement improved precision without reducing recall, indicating enhanced signal clarity through measurement-level stabilization. Explainable AI analysis further revealed that word reading and reading fluency contributed strongly to model-level decision boundaries, while vocabulary knowledge provided complementary influence at the individual level. These findings provide prototype-level evidence that a human-validated, multimodal screening workflow can support stable early-risk detection. From a sustainability perspective, the results suggest potential design-level contributions to improving accessibility and reducing delays in early identification processes.