Development and validation of SARA: an AI-assisted reading assessment tool for primary school children in low-resource Spanish-speaking contexts
Daniel Cubilla-Bonnetier, Enzo Casamassima, Hugo Marte-Santana, Laura V. Sánchez-Vincitore
Low reading achievement is a persistent challenge in Latin American education, particularly in the Dominican Republic, where most primary school students perform below minimum proficiency levels. Standardized, culturally adapted assessment tools are scarce in Spanish-speaking low-resource contexts, limiting early identification of reading difficulties at scale. This paper presents the development and validation of SARA (Self-Applied Reading Assessment), an AI-assisted tablet-based tool for autonomous administration and automated scoring of reading assessments in Dominican primary school children (grades 2–6). SARA integrates ECLEC (Evaluación Centroamericana de la Lectura), a battery grounded in the Simple View of Reading and the Dual-Route Model, covering decoding (letter identification, syllable, word, and pseudoword reading) and oral and reading comprehension. Following convergent validity testing against the PROLEC-R and two pilot studies, SARA was administered to 1,860 students from randomly selected public and private schools in Santo Domingo. Oral responses were scored by two independent human rater teams, yielding an overall inter-rater agreement of 86.11% (73.07%–94.70% across subtests) and a Cohen's Kappa of.550, a conservative estimate reflecting the class imbalance inherent in reading assessment data rather than poor rater reliability. Machine learning models were trained using the expert rater team's labels and evaluated against a conservative human consensus, defined by cases in which both teams agreed. Results showed good to excellent internal consistency across decoding subtests and adequate reliability for reading comprehension. Structural equation models confirmed construct validity by replicating the predicted developmental shift from decoding ( β = .70 in 2nd grade) to oral comprehension ( β = .60 in 6th grade) as predictors of reading comprehension. On the held-out test set, AI scoring achieved 91.97% overall accuracy, comparable to the 86.11% agreement between human raters, with class-specific accuracy higher for correct (90.0%–96.2%) than incorrect responses (73.9%–80.1%). Grade- and school-system-differentiated norms were established, reflecting the documented performance gap between public and private schools. SARA shows promising initial evidence of validity and reliability as a reading assessment tool for Spanish-speaking low-resource contexts and may constitute a scalable solution pending further validation in other regions and contexts.