• AI IN LEARNING ASSESSMENT: REDUCING BIAS OR REPRODUCING INEQUALITIES? A COMPARATIVE STUDY OF HUMAN AND AI SCORING IN FRENCH AS A FOREIGN LANGUAGE
  •  
  • Hajar SALIM
  • ENS/ Moulay Ismail University, Meknes ,Morocco
  • ORCID iD :0009-0008-5122-1571
  • haj.salim@edu.umi.ac.ma
  • &
  • Awatif BEGGAR
  • FS/Moulay Ismail University, Meknes, Morocco
  • ORCID iD :0009-0002-6500-3096
  • awatifbeggar@gmail.com

Introduction: The gradual integration of artificial intelligence into educational processes continues to intensify, with the aim of enriching the learning experience at all levels. Researchers, teachers, and experts in educational technology are striving to keep pace with this rapid transformation by actively participating in digital transformation workshops. Seeking to identify concrete and effective ways to innovate and integrate these innovations into teaching practices. This objective also fits within a context marked by international linguistic openness, particularly with reference to proficiency in French as a foreign language. However, the assessment component seems to remain stuck in traditional patterns. Often, it  reduced to a standardized, linear grading system that is sometimes unrelated to the specific needs of learners and their learning contexts. This model can, in some cases, contribute to academic failure or dropout. However, educational assessment cannot be limited to a simple measurement of what has been learned. It remains an essential component of the educational process, guiding the direction of learning, identifying learners’ needs, and adapting teaching strategies. This is all the more important in language teaching, which currently relies as much on the development of communicative and cultural competencies as on knowledge of linguistic rules.

In language teaching, the assessment of skills-whether oral or written, receptive or productive-poses a genuine pedagogical challenge. It requires rigorous, personalized approaches tailored to learners’ profiles. Yet this discipline is often practiced without any real questioning of its foundations or any desire for improvement. Current assessment systems suffer from numerous biases and limitations linked to factors that call into question the objectivity, validity, and reliability of the assessment process. These observations have led us to examine existing mechanisms and explore new avenues for rethinking assessment practices. Can we envisage AI providing concrete solutions to these challenges? Could it help to objectify and standardize the assessment of language skills, or, on the contrary, does it risk introducing new biases linked to its algorithms and the data on which it relies? Thus, this research revolves around the central question: To what extent can AI, as an assessment tool, address the limitations of the traditional system without creating new ones? Does AI promise automated, rapid, and standardized grading that could reduce certain forms of subjectivity, or does it raise questions about the transparency of algorithms and the quality of the data? Guided by the research questions, the following hypotheses are proposed: (1) AI-assisted grading helps reduce certain biases associated with human judgment; (2) Despite these advantages, it has limitations in evaluating written work, particularly when it comes to assessing linguistic nuances, context, and creative elements. With this in perspective, the article is organized into four sections, which contribute to the examination of the issue and the testing of the research hypotheses. We first present the theoretical framework of the study, before describing the methodology adopted. The results are then presented, analyzed, and subsequently discussed in light of the literature to highlight the main contributions of the research.

Abstract: Education has not been left behind in the face of change; as new technologies have significantly transformed the educational landscape. The integration of digital technology into teaching practices has opened the door to more diverse learning methods, greater access to resources, and more flexible and personalized instruction, creating a forum for discussing how to leverage artificial intelligence as a catalyst to enrich the educational experience. Assessment, as an essential component of the teaching-learning process, often remains on the sidelines of these developments. It is our responsibility to carefully plan its design, administration, grading, and analysis in order to fully leverage its central role in language learning and minimize its well-documented biases. This article analyzes whether automated assessment-and more specifically, assessment using AI-can correct these biases or whether, on the contrary, it risks introducing new ones. The study is based on an experiment conducted with 90 Higher Education Cycle (Semester 1), whose exams were graded by two teachers and an AI grader. Significant discrepancies emerged. While the AI reduces some of these biases and brings standardization and consistency to grading, algorithmic biases and a lack of consideration for qualitative dimensions also emerged. Consequently, a hybrid approach combining AI and human teachers is recommended as a balanced solution.

Keywords : Educational assessment; Artificial intelligence ; Assessment bias ; ChatGPT5

References

  • Bressoux, P., Dompnier, B., & Pansu, P. (2011). School Assessment: A Multifactorial Activity. In Assessment: A Threat?
  • Bressoux, P., & Pansu, P. (2003). When Teachers Assess Their Students. Presses universitaires de France.
  • Bulut, O., et al. (2024). The Rise of Artificial Intelligence in Educational Measurement: Opportunities and Ethical Challenges. arXiv. 10.48550/arXiv.2406.18900
  • CNESCO. (2023). School Assessment and Grading: Expert Reports. Paris: National Council for the Evaluation of the School System. https://www.cnesco.fr/wp-content/uploads/2024/07/Cnesco-CC-Eval_Notes-des-experts.pdf
  • De Ketele, J.-M., & Gerard, F.-M. (2005). Validation of Assessment Tests Using a Competency-Based Approach. Measurement and Evaluation in Education, 28(3), 1.
  • De Ketele, J.-M. (2011). Assessment and the Curriculum: Conceptual Foundations, Debates, and Issues. Educational Science Files, 25, 89–106.
  • Gobrecht et al. (2024). Beyond human subjectivity and error: A novel AI grading system. arXiv.
  • Gong, W. (2022). Reshaping the EFL Formative Assessment Pedagogy with Blockchain Technology. International Journal of English Linguistics, 13(1), 12. https://doi.org/10.5539/ijel.v13n1p12
  •