CORTEXA
← Browse
crossrefAI2026-07-18Cited by 0

Predicting Student Stress Using Machine Learning Ensemble Models: A Multi-Criteria Comparison with Explainable Artificial Intelligence Analysis

Daniel Cristóbal Andrade-Girón, William Joel Marin-Rodriguez, Marcelo Gumercindo Zuñiga-Rojas, Abrahan Cesar Neri-Ayala, Edgar Tito Susanibar-Ramírez, Miguel Angel Aguilar-Luna-Victoria

Student stress is a significant mental health issue in educational settings; therefore, developing reliable, calibrated, and interpretable predictive models can support the classification of observed stress levels. This study analyzed the public Student Stress Factors dataset, comprising 1100 records, 20 predictors, and one target variable, using a supervised machine learning pipeline designed to reduce information leakage. The pipeline included stratified data partitioning, encapsulated preprocessing, nested cross-validation restricted to the training data, and independent holdout evaluation. Nine ensemble and boosting algorithms for tabular data were compared: AdaBoost, Gradient Boosting, Random Forest, Extra Trees, Bagging, Voting, Stacking, XGBoost, and LightGBM. Model performance was assessed using key discrimination and calibration metrics, together with the nonparametric Friedman test for statistical comparison. Gradient Boosting achieved the best average performance in nested cross-validation, with an accuracy of 89.55 ± 3.16%, F1-weighted of 89.54 ± 3.17%, MCC of 0.845 ± 0.047, and ROC-AUC weighted of 98.59 ± 0.92%. XGBoost and LightGBM showed comparable performance. In the independent holdout set, the final calibrated model maintained robust predictive performance, achieving an accuracy of 0.8818, F1-weighted of 0.8818, MCC of 0.8237, and ROC-AUC weighted of 0.9861. Although the overall results indicate stable and high predictive performance, the Friedman test did not identify statistically significant differences among the algorithms, χ2 = 10.953, p = 0.204. Therefore, model selection should consider not only predictive accuracy but also computational efficiency, calibration, interpretability, and implementation feasibility. Despite the internal stability of the pipeline and satisfactory holdout performance, the public and cross-sectional nature of the dataset limits causal inference and model transferability. Consequently, external and prospective validation is required before integration into institutional early warning systems.

View free PDFSource page

Related papers

crossrefAI2026-02-01Cited by 3

Enhancing Decision Intelligence Using Hybrid Machine Learning Framework with Linear Programming for Enterprise Project Selection and Portfolio Optimization

Abdullah, Nida Hafeez, Carlos Guzmán Sánchez-Mejorada, Miguel Jesús Torres Ruiz, Rolando Quintero Téllez, Eponon Anvi Alex, et al.

This study presents a hybrid analytical framework that enhances project selection by achieving reasonable predictive accuracy through the integration of expert judgment and modern artificial intelligence (AI) techniques. Using an enterprise-level dataset of 10,000 completed softw…

View free PDFSource page
crossrefAI2026-05-09

Machine Learning Models for Predicting Post-Hepatectomy Liver Failure: A Systematic Review

Calin Muntean, Vasile Gaborean, Razvan Constantin Vonica, Sebastian Aurelian Stefaniga, Alaviana Monique Faur, Catalin Vladut Ionut Feier

Background and Objectives: Post-hepatectomy liver failure (PHLF) remains the leading cause of mortality following hepatic resection, with reported incidence rates ranging from 1.2% to 32%. Traditional scoring systems such as the Child–Pugh score, Model for End-Stage Liver Disease…

View free PDFSource page
crossrefAI2026-07-01

Towards Data-Driven Weather Intelligence in Palestine: A Multi-Station Benchmark of Classical Machine Learning and Deep Learning Models

Mohammad Odeh, Ahmad Hasasneh

Precise weather forecasting plays a critical role in sectors such as agriculture, transport, energy management, and climate change adaptation, and machine learning and deep learning algorithms have been widely used for data-driven time series forecasting problems. In this work, w…

View free PDFSource page
crossrefAI2026-06-01

Beyond Vital Signs: A Machine Learning Model Using Comprehensive Triage-Time Data to Detect Undertriage in Emergency Department Patients

Kyungman Cha, Sohee Lee, Jaekwang Shin, Jee Yong Lim

Undertriage—the misclassification of acutely ill patients into low-acuity triage categories—is a persistent patient safety concern, and prior machine learning approaches restricted to vital signs have yielded modest predictive performance. We hypothesized that this ceiling reflec…

View free PDFSource page
crossrefAI2026-07-12

eGFR-AI: A Stacked Machine-Learning Model for Early Postoperative Kidney Function Prediction—A Pilot Study

Eva Brenner, Luka Bulić, Vilena Vrbanović Mijatović

Background: Postoperative kidney dysfunction is a common and serious complication in surgical patients. Kidney function is typically assessed using the estimated glomerular filtration rate (eGFR), most often calculated with the CKD-EPI equation based on serum creatinine. While se…

View free PDFSource page
crossrefAI2026-01-25Cited by 4

A Hybrid Intrusion Detection Framework Using Deep Autoencoder and Machine Learning Models

Salam Allawi Hussein, Sándor R. Répás

This study provides a detailed comparative analysis of a three-hybrid intrusion detection method aimed at strengthening network security through precise and adaptive threat identification. The proposed framework integrates an Autoencoder-Gaussian Mixture Model (AE-GMM) with two s…

View free PDFSource page