• CROSS-UNIVERSITY GENERALIZABILITY AND EXPLAINABILITY OF ACADEMIC RISK PREDICTION IN KINSHASA HIGHER EDUCATION
  •  
  • Augustin PAMBI TADIAMBA
  • University of Kinshasa, Democratic Republic of the Congo
  • ORCID iD: 0009-0007-3284-0660
  • augustintadiamba1@gmail.com

Introduction: Universities increasingly use educational data to move from retrospective reporting toward earlier forms of academic support. Grades, attendance records, assignment submissions, prior credits, and other traces can be assembled while a semester is still under way, giving teachers and academic services an opportunity to act before difficulties become irreversible (Long & Siemens, 2011; Romero & Ventura, 2020). However, the predictive performance achieved in a consolidated data set does not guarantee that the model would hold its accuracy when implemented in a new institution. A classifier can find patterns that are robust across different contexts, but it can also exploit local grading standards, program structures, or data collection methods that are not generalizable. This distinction matters in Kinshasa. Universities belong to the same national system of higher education, but vary in size, range of programs, number of students, administrative procedures, and degree of organization applied to academic knowledge.A model that looks satisfactory when all observations are mixed may hide uneven performance by institution. The practical question is therefore no longer only whether academic risk can be predicted, but whether the predictive relationship survives a change of university and whether the same variables continue to carry information in the new setting. A recent study conducted in Kinshasa showed that six-level academic-risk prediction can be developed from mid-semester educational records and linked to graduated support options (Tadiamba et al., in press). An important question nevertheless remains unresolved: whether the learned relationships remain stable when the model is transferred from one university to another. The present study addresses this gap by treating each institution as a separate validation domain and examining external performance, directional transfer, calibration, class-level errors, and institution-specific permutation importance. The research question is: to what extent does a fixed academic-risk model preserve discrimination, class balance, probability reliability, and explanatory structure when it is transferred among the University of Kinshasa (UNIKIN), National Pedagogical University (UPN), and the Protestant University of Congo (UPC)? Three objectives follow from this question. The first is to compare internal performance with performance obtained when an entire university is excluded from training. The second is to determine whether the dominant predictors are shared or institution-dependent. The third is to identify the classes and transfer directions in which errors become more pronounced. The approach relies on three operational hypotheses. The external macro F1 score for each institution will not deviate by more than 0.05 from the within-university cross-validation score. H2 indicates that ongoing evaluation, midterm performance, and already validated credits will continue to be the primary predictors across all three universities. The extreme categories of the ordinal scale (Risk and Excellence) will exhibit greater recall instability than the intermediate categories. These hypotheses are evaluated as empirical expectations rather than causal assertions. The contribution is institutional as well as methodological. The study methodologically transitions the unit of validation from an individual randomly selected student record to an entire institution, representing a more stringent assessment of portability. Institutionally, it provides evidence to determine if the universities of Kinshasa can initiate a unified predictive core or if entirely distinct models are required from the outset. The paper thereafter addresses methodology, results and analysis, as well as limitations and discussion.

Abstract: This paper studies the predictive performance and transferability of an academic-risk prediction model for three Kinshasa institutions. We employ a standardized data collection of 9,000 student-semester observations from the Protestant University of Congo, the National Pedagogical University and the University of Kinshasa. We consider institutional generalizability, and use a common random-forest specification across the investigations. University membership is not part of the predictor matrix but is used to define the training and validation domains. Thirteen academic, historical, behavioral and environmental aspects are examined with a repeatable Python/Jupyter Notebook approach. We evaluate the top-label calibration, analyze the class-level errors, analyze the permutation importance, test the directional transfer between institutions, perform leave-one-university-out external validation, and perform intra-university cross-validation. The external macro F1 ratings of UNIKIN, UPN and UPC are 0.658, 0.706 and 0.718 respectively. The macro one-vs-rest AUC ranged from 0.944 to 0.949. The three major predictors, in each institution, were the continuous assessment average, the midterm average and the previously verified credit rate, but their relative weights differed. The major weaknesses were found in the extreme poles of the 6-level performance scale, especially with UNIKIN’s Excellence. The findings do not support unqualified implementation, but they do support a shared predictive core across the three institutional contexts . The model needs to be periodically recalibrated (even for probabilities) and externally validated for a local context before it can be used in academic decision-support systems, and its performance in the minority class needs to be evaluated.

Keywords: academic risk prediction, random forest, explainability, cross-university generalizability, external validation.

References

  • Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324
  • Kaufman, S., Rosset, S., Perlich, C., & Stitelman, O. (2012). Leakage in data mining: Formulation, detection, and avoidance. ACM Transactions on Knowledge Discovery from Data, 6(4), Article 15. https://doi.org/10.1145/2382577.2382579
  • Long, P., & Siemens, G. (2011). Penetrating the fog: Analytics in learning and education. EDUCAUSE Review, 46(5), 30–40.
  • Romero, C., & Ventura, S. (2020). Educational data mining and learning analytics: An updated survey. WIREs Data Mining and Knowledge Discovery, 10(3), e1355. https://doi.org/10.1002/widm.1355
  • Slade, S., & Prinsloo, P. (2013). Learning analytics: Ethical issues and dilemmas. American Behavioral Scientist, 57(10), 1510–1529. https://doi.org/10.1177/0002764213479366
  • Tadiamba, A. P., Kafunda, P. K., Kutangila, D. M., Minga, J. S., Ngwaba, B. B., Ngwaba, J. B., Tshimanga, G.-R. G., Katshitshi, J.-J. M., & Mboma, R. K. (in press). Academic risk profiling and tailored student guidance through machine learning in Kinshasa universities. International Journal of Innovative Science and Research Technology.
  • UNESCO. (2021). Recommendation on the ethics of artificial intelligence. UNESCO.