Design and Development of a Transparent AI-Driven Assessment Framework Based on BERT/RoBERTa and Explainable AI for Student Resume Evaluation

Yogi Prasetyo, Yuli Sopianti, Latifahny Aridia Alfitri

Abstract


The evaluation of resume-based answers demands complex semantic understanding and knowledge structuring, making manual assessment prone to inter-rater inconsistency, while deep learning models such as BERT generally operate as a black box without adequate transparency. This study aims to design a Transparent AI-Driven Assessment Framework that integrates a BERT/RoBERTa model for resume assessment, an Explainable AI module based on SHAP and LIME, and a dual-validation mechanism based on multiple-choice analytics to verify learners' conceptual understanding. The study employs the Research and Development (R&D) method, covering needs analysis, architectural design, and the development of a prototype based on a microservices architecture interoperable with the Learning Management System. A formative evaluation of the prototype conducted on 15 respondents (7 lecturers and 8 students) yielded an average System Usability Scale score of 88.17 (Excellent category) and high perceived usefulness and transparency scores (4.78 and 4.72 out of 5), which preliminarily support the feasibility of the design, although these results remain indicative given the limited number of respondents.


Keywords


Artificial Intelligence; Explainable Artificial Intelligence; Dual-validation Mechanism; Intelligent Assessment; Learning Analytics.

Full Text:

PDF

References


Bangor, A., Kortum, P., & Miller, J. (2009). Determining what individual SUS scores mean: Adding an adjective rating scale. Journal of Usability Studies, 4(3), 114–123.

Bazza, H., Bimonte, S., Gourti, Z., Rizzi, S., & Badir, H. (2025). An MDA approach for robotic-based real-time business intelligence applications. Data and Knowledge Engineering, 157, 102418. https://doi.org/10.1016/j.datak.2025.102418

Beaumont, C., O'Doherty, M., & Shannon, L. (2011). Reconceptualising assessment as a weapon for learning. European Journal of Engineering Education, 36(2), 137-147. https://doi.org/10.1080/03043797.2010.547941

Brooke, J. (1996). SUS: A quick and dirty usability scale. In P. W. Jordan, B. Thomas, B. A. Weerdmeester, & A. L. McClelland (Eds.), Usability evaluation in industry (pp. 189–194). Taylor and Francis.

Carless, D., & Boud, D. (2018). Student voice and learning: Conceptualising capabilities for dialogic engagement. Studies in Higher Education, 43(9), 1568-1578. https://doi.org/10.1080/03075079.2016.1265196

Chaikhamwang, S., Montri, W., Fongmanee, S., & Janthajirakowit, C. (2025). Application of advanced natural language processing methodologies utilizing GPT-4o-mini to augment the efficacy of an essay examination framework on a web-based application. International Journal of Computer Applications, 6(1), 1–30. https://doi.org/10.34218/IJCA_06_01_001

Conijn, R., Kahr, P., & Snijders, C. (2023). The effects of explanations in automated essay scoring systems on student trust and motivation. Journal of Learning Analytics, 10(1), 37–53. https://doi.org/10.18608/jla.2023.7801

Davis, F. D. (1989). Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Quarterly, 13(3), 319–340. https://doi.org/10.2307/249008

Emirtekin, E. (2025). Large language model-powered automated assessment: A systematic review. Applied Sciences, 15(10), 5683. https://doi.org/10.3390/app15105683

Guskey, T. R. (2024). Addressing inconsistencies in grading practices. Phi Delta Kappan, 105(8), 52–57. https://doi.org/10.1177/00317217241251883

Hattie & Timperley. (2007). Bukti bahwa timely, specific, actionable feedback meningkatkan learning outcomes

Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81-112. https://doi.org/10.3102/003465430298487

Hermosilla, P., Berrios, S., & Allende-Cid, H. (2025). Explainable AI for forensic analysis: A comparative study of SHAP and LIME in intrusion detection models. Applied Sciences, 15(13), 7329. https://doi.org/10.3390/app15137329

Ilieva, G., Yankova, T., Ruseva, M., & Kabaivanov, S. (2025). A framework for generative AI-driven assessment in higher education. Information, 16(6), 472. https://doi.org/10.3390/info16060472

Kates, S., Paulsen, T., Yntiso, S., & Tucker, J. A. (2023). Bridging the grade gap: Reducing assessment bias in a multi-grader class. Political Analysis, 31(4), 642–650. https://doi.org/10.1017/pan.2022.27

Linder, A., Nordin, M., Gerdtham, U., & Heckley, G. (2023). Grading bias and young adult mental health. Health Economics, 32(3), 675–696. https://doi.org/10.1002/hec.4639

Malouff, J. M., & Thorsteinsson, E. B. (2016). Bias in grading: A meta-analysis of experimental research findings. Australian Journal of Education, 60(3), 245–256. https://doi.org/10.1177/0004944116664618

Mpolomoka, D. L. (2025). Utilizing artificial intelligence for assessment in higher education. Pedagogical Research, 10(3), em0243. https://doi.org/10.29333/pr/16677

Prasetyo, Y. (2021). Perencanaan arsitektur enterprise smart school menggunakan TOGAF: Studi kasus SMK Negeri 13 Bandung.

Rani, D., & Badusha, B. (2023). Intelligent descriptive answer evaluation system. Computer Science, Engineering and Technology, 1(3), 13–16. https://doi.org/10.46632/cset/1/3/3

Rintala, A. (2023). Assessment analysis: Methods and implementation options for multiple-choice exams. International Journal of Innovation in Education, 8(1), 20. https://doi.org/10.1504/IJIIE.2023.128460

Salih, A. M., et al. (2025). A perspective on explainable artificial intelligence methods: SHAP and LIME. Advanced Intelligent Systems, 7(1), 2400304. https://doi.org/10.1002/aisy.202400304

UNESCO. (2023). Artificial intelligence and education: Guidance on generative AI in education and research. UNESCO Publishing.

Wanichsan, D., Panjaburee, P., & Chookaew, S. (2021). Enhancing knowledge integration from multiple experts to guiding personalized learning paths for testing and diagnostic systems. Computers and Education: Artificial Intelligence, 2, 100013. https://doi.org/10.1016/j.caeai.2021.100013

Zhang, G., Pan, L., Tang, F., & Yao, F. (2025). Explainable artificial intelligence in the talent recruitment process—a literature review. Cogent Business and Management, 12(1), 2570881. https://doi.org/10.1080/23311975.2025.2570881




DOI: https://doi.org/10.31764/justek.v9i3.41616

Refbacks

  • There are currently no refbacks.


JUSTEK : Jurnal Sains dan Teknologi sudah terindeks

EDITORIAL OFFICE: