Design and Development of a Transparent AI-Driven Assessment Framework Based on BERT/RoBERTa and Explainable AI for Student Resume Evaluation
Abstract
The evaluation of resume-based answers demands complex semantic understanding and knowledge structuring, making manual assessment prone to inter-rater inconsistency, while deep learning models such as BERT generally operate as a black box without adequate transparency. This study aims to design a Transparent AI-Driven Assessment Framework that integrates a BERT/RoBERTa model for resume assessment, an Explainable AI module based on SHAP and LIME, and a dual-validation mechanism based on multiple-choice analytics to verify learners' conceptual understanding. The study employs the Research and Development (R&D) method, covering needs analysis, architectural design, and the development of a prototype based on a microservices architecture interoperable with the Learning Management System. A formative evaluation of the prototype conducted on 15 respondents (7 lecturers and 8 students) yielded an average System Usability Scale score of 88.17 (Excellent category) and high perceived usefulness and transparency scores (4.78 and 4.72 out of 5), which preliminarily support the feasibility of the design, although these results remain indicative given the limited number of respondents.
Keywords
Full Text:
PDFReferences
Bangor, A., Kortum, P., & Miller, J. (2009). Determining what individual SUS scores mean: Adding an adjective rating scale. Journal of Usability Studies, 4(3), 114–123.
Bazza, H., Bimonte, S., Gourti, Z., Rizzi, S., & Badir, H. (2025). An MDA approach for robotic-based real-time business intelligence applications. Data and Knowledge Engineering, 157, 102418. https://doi.org/10.1016/j.datak.2025.102418
Beaumont, C., O'Doherty, M., & Shannon, L. (2011). Reconceptualising assessment as a weapon for learning. European Journal of Engineering Education, 36(2), 137-147. https://doi.org/10.1080/03043797.2010.547941
Brooke, J. (1996). SUS: A quick and dirty usability scale. In P. W. Jordan, B. Thomas, B. A. Weerdmeester, & A. L. McClelland (Eds.), Usability evaluation in industry (pp. 189–194). Taylor and Francis.
Carless, D., & Boud, D. (2018). Student voice and learning: Conceptualising capabilities for dialogic engagement. Studies in Higher Education, 43(9), 1568-1578. https://doi.org/10.1080/03075079.2016.1265196
Chaikhamwang, S., Montri, W., Fongmanee, S., & Janthajirakowit, C. (2025). Application of advanced natural language processing methodologies utilizing GPT-4o-mini to augment the efficacy of an essay examination framework on a web-based application. International Journal of Computer Applications, 6(1), 1–30. https://doi.org/10.34218/IJCA_06_01_001
Conijn, R., Kahr, P., & Snijders, C. (2023). The effects of explanations in automated essay scoring systems on student trust and motivation. Journal of Learning Analytics, 10(1), 37–53. https://doi.org/10.18608/jla.2023.7801
Davis, F. D. (1989). Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Quarterly, 13(3), 319–340. https://doi.org/10.2307/249008
Emirtekin, E. (2025). Large language model-powered automated assessment: A systematic review. Applied Sciences, 15(10), 5683. https://doi.org/10.3390/app15105683
Guskey, T. R. (2024). Addressing inconsistencies in grading practices. Phi Delta Kappan, 105(8), 52–57. https://doi.org/10.1177/00317217241251883
Hattie & Timperley. (2007). Bukti bahwa timely, specific, actionable feedback meningkatkan learning outcomes
Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81-112. https://doi.org/10.3102/003465430298487
Hermosilla, P., Berrios, S., & Allende-Cid, H. (2025). Explainable AI for forensic analysis: A comparative study of SHAP and LIME in intrusion detection models. Applied Sciences, 15(13), 7329. https://doi.org/10.3390/app15137329
Ilieva, G., Yankova, T., Ruseva, M., & Kabaivanov, S. (2025). A framework for generative AI-driven assessment in higher education. Information, 16(6), 472. https://doi.org/10.3390/info16060472
Kates, S., Paulsen, T., Yntiso, S., & Tucker, J. A. (2023). Bridging the grade gap: Reducing assessment bias in a multi-grader class. Political Analysis, 31(4), 642–650. https://doi.org/10.1017/pan.2022.27
Linder, A., Nordin, M., Gerdtham, U., & Heckley, G. (2023). Grading bias and young adult mental health. Health Economics, 32(3), 675–696. https://doi.org/10.1002/hec.4639
Malouff, J. M., & Thorsteinsson, E. B. (2016). Bias in grading: A meta-analysis of experimental research findings. Australian Journal of Education, 60(3), 245–256. https://doi.org/10.1177/0004944116664618
Mpolomoka, D. L. (2025). Utilizing artificial intelligence for assessment in higher education. Pedagogical Research, 10(3), em0243. https://doi.org/10.29333/pr/16677
Prasetyo, Y. (2021). Perencanaan arsitektur enterprise smart school menggunakan TOGAF: Studi kasus SMK Negeri 13 Bandung.
Rani, D., & Badusha, B. (2023). Intelligent descriptive answer evaluation system. Computer Science, Engineering and Technology, 1(3), 13–16. https://doi.org/10.46632/cset/1/3/3
Rintala, A. (2023). Assessment analysis: Methods and implementation options for multiple-choice exams. International Journal of Innovation in Education, 8(1), 20. https://doi.org/10.1504/IJIIE.2023.128460
Salih, A. M., et al. (2025). A perspective on explainable artificial intelligence methods: SHAP and LIME. Advanced Intelligent Systems, 7(1), 2400304. https://doi.org/10.1002/aisy.202400304
UNESCO. (2023). Artificial intelligence and education: Guidance on generative AI in education and research. UNESCO Publishing.
Wanichsan, D., Panjaburee, P., & Chookaew, S. (2021). Enhancing knowledge integration from multiple experts to guiding personalized learning paths for testing and diagnostic systems. Computers and Education: Artificial Intelligence, 2, 100013. https://doi.org/10.1016/j.caeai.2021.100013
Zhang, G., Pan, L., Tang, F., & Yao, F. (2025). Explainable artificial intelligence in the talent recruitment process—a literature review. Cogent Business and Management, 12(1), 2570881. https://doi.org/10.1080/23311975.2025.2570881
DOI: https://doi.org/10.31764/justek.v9i3.41616
Refbacks
- There are currently no refbacks.
JUSTEK : Jurnal Sains dan Teknologi sudah terindeks
EDITORIAL OFFICE:









