Scalability Analysis and Computational Performance of BiLSTM-CRF Model in Indonesian Named Entity Recognition

Nurul Isnaeni Rahmat, Emha Taufiq Luthfi

Abstract


Named Entity Recognition (NER) is a fundamental task in natural language processing that supports information extraction and knowledge organization. However, empirical studies examining the computational scalability of conventional NER models for the Indonesian language remain limited. This study investigates the scalability and computational performance of the BiLSTM-CRF model for Indonesian NER tasks. The objective is to evaluate how the model’s computational requirements and predictive performance change as the size of the training dataset increases. An experimental evaluation was conducted by training the BiLSTM-CRF model on three dataset scales derived from the WikiANN corpus (small, medium, and large) using a standard configuration with randomly initialized embeddings in a CPU-based environment. Model performance was assessed using the F1-score, while computational scalability was analyzed through measurements of training time, memory consumption, and inference speed. The results indicate a clear scalability pattern in which computational costs increase with dataset size, particularly in training time and memory usage. At the same time, predictive performance improves as more training data becomes available, with the F1-score increasing from 0.70 on the smallest dataset to 0.86 on the largest dataset. These findings provide empirical evidence on the scalability behavior of the BiLSTM-CRF model for Indonesian NER and offer practical insights for selecting model configurations under limited computational resources.

 


Keywords


BiLSTM-CRF; Named Entity Recognition; Model Scalability; CPU Computation; Indonesian Language.

Full Text:

DOWNLOAD [PDF]

References


Ahmed, S. F., Alam, M. S. Bin, Hassan, M., Rozbu, M. R., Ishtiak, T., Rafa, N., Mofijur, M., Shawkat Ali, A. B. M., & Gandomi, A. H. (2023). Deep learning modelling techniques: current progress, applications, advantages, and challenges. Artificial Intelligence Review, 56(11), 13521–13617. https://doi.org/10.1007/s10462-023-10466-8

Aljumaily, H., Laefer, D. F., Cuadra, D., & Velasco, M. (2023). Point cloud voxel classification of aerial urban LiDAR using voxel attributes and random forest approach. In International Journal of Applied Earth Observation and Geoinformation (Vol. 118). Elsevier B.V. https://doi.org/10.1016/j.jag.2023.103208

Ashebir, D., & Tadesse, G. (2022). Named Entity Recognition for Hadiyya Language using BiLSTM-CRF Model. Indian Journal Of Science And Technology, 15(47), 2612–2618. https://doi.org/10.17485/IJST/v15i47.1090

Budi, I., & Suryono, R. R. (2023). Application of named entity recognition method for Indonesian datasets: a review. In Bulletin of Electrical Engineering and Informatics (Vol. 12, Number 2, pp. 969–978). Institute of Advanced Engineering and Science. https://doi.org/10.11591/eei.v12i2.4529

Chen, C., Zhang, P., Zhang, H., Dai, J., Yi, Y., Zhang, H., Zhang, Y., & Khan, M. J. (2020). Deep Learning on Computational-Resource-Limited Platforms: A Survey. Mobile Information Systems, (Article ID 8454327), 1–19. https://doi.org/10.1155/2020/8454327

Dagdelen, J., Dunn, A., Lee, S., Walker, N., Rosen, A. S., Ceder, G., Persson, K. A., & Jain, A. (2024). Structured information extraction from scientific text with large language models. Nature Communications, 15(1), 1–14. https://doi.org/10.1038/s41467-024-45563-x

Gayathri, C., & Samson Ravindran, D. R. (2025). Named entity recognition using Bi-LSTM model with pointer cascade conditional random field for selecting high-profit products. Egyptian Informatics Journal, 31(100703), 1–14. https://doi.org/10.1016/j.eij.2025.100703

Guntreddi, V., & V, S. (2025). Deep learning based glaucoma detection using majority voting ensemble of ResNet50, VGG16, and Swin Transformer. Results in Engineering, 28(107229), 1–13. https://doi.org/10.1016/j.rineng.2025.107229

Hafsa, N. E., Alzoubi, H. M., & Almutiq, A. S. (2025). Accurate disaster entity recognition based on contextual embeddings in self-attentive BiLSTM-CRF. Plos One, 20(3), 1–27. https://doi.org/10.1371/journal.pone

Jain, A., Kulkarni, G., & Shah, V. (2018). Natural Language Processing. International Journal of Computer Sciences and Engineering, 6(1), 161–167. https://doi.org/10.26438/ijcse/v6i1.161167

Jehangir, B., Radhakrishnan, S., & Agarwal, R. (2023). A survey on Named Entity Recognition — datasets, tools, and methodologies. Natural Language Processing Journal, 3, 100017. https://doi.org/10.1016/j.nlp.2023.100017

Keraghel, I., Morbieu, S., & Nadif, M. (2024). Recent Advances in Named Entity Recognition: A Comprehensive Survey and Comparative Study. In International Conference on Artificial Neural Networks (ICANN), 1–42. https://doi.org/http://arxiv.org/abs/2401.10825

Khurana, D., Koli, A., Khatter, K., & Singh, S. (2023). Natural language processing: state of the art, current trends and challenges. Multimedia Tools and Applications, 82(3), 3713–3744. https://doi.org/10.1007/s11042-022-13428-4

Kumar, A., Sharma, R., & Bedi, P. (2024). Towards Optimal NLP Solutions: Analyzing GPT and LLaMA-2 Models Across Model Scale, Dataset Size, and Task Diversity. Engineering, Technology and Applied Science Research, 14(3), 14219–14224. https://doi.org/10.48084/etasr.7200

Kusumawardani, R. P., & Kusumawati, K. N. (2024). Named entity recognition in the medical domain for Indonesian language health consultation services using bidirectional-lstmcrf algorithm. 9th International Conference on Computer Science and Computational Intelligence 2024 (ICCSCI 2024), 245, 1146–1156. https://doi.org/10.1016/j.procs.2024.10.344

Li, J., Sun, A., Han, J., & Li, C. (2020). A Survey on Deep Learning for Named Entity Recognition. http://neuroner.com/

Li, Q., Peng, H., Li, J., Xia, C., Yang, R., Sun, L., Yu, P. S., & He, L. (2022). A Survey on Text Classification: From Traditional to Deep Learning. In ACM Transactions on Intelligent Systems and Technology (Vol. 13, Number 2). Association for Computing Machinery. https://doi.org/10.1145/3495162

Ma, P., Jiang, B., Lu, Z., Li, N., & Jiang, Z. (2021). Cybersecurity Named Entity Recognition Using Bidirectional Long Short-Term Memory with Conditional Random Fields. Tsinghua Science and Technology, 26(3), 259–265. https://doi.org/10.26599/TST.2019.9010033

Marreddy, M., Oota, S. R., Vakada, L. S., Chinni, V. C., & Mamidi, R. (2022). Am I a Resource-Poor Language? Data Sets, Embeddings, Models and Analysis for four different NLP Tasks in Telugu Language. ACM Transactions on Asian and Low-Resource Language Information Processing, 22(1), 1–35. https://doi.org/10.1145/3531535

Menghani, G. (2023). Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better. In ACM Computing Surveys (Vol. 55, Number 12). Association for Computing Machinery. https://doi.org/10.1145/3578938

Minaee, S., Kalchbrenner, N., Cambria, E., Nikzad, N., Chenaghlu, M., & Gao, J. (2022). Deep Learning-Based Text Classification. ACM Computing Surveys, 54(3), 1–40. https://doi.org/10.1145/3439726

Mohamed, A., Najafabadi, M. K., Wah, Y. B., Zaman, E. A. K., & Maskat, R. (2020). The state of the art and taxonomy of big data analytics: view from new big data framework. Artificial Intelligence Review, 53(2), 989–1037. https://doi.org/10.1007/s10462-019-09685-9

Murakami, E., Shionoya, T., Komenoi, S., Suzuki, Y., & Sakane, F. (2025). Cloning and characterization of novel testis-Specific diacylglycerol kinase η splice variants 3 and 4. PLoS ONE, 11(9), 1–14. https://doi.org/10.1371/journal.pone

Nabiilah, G. Z., Alam, I. N., Purwanto, E. S., & Hidayat, M. F. (2024). Indonesian multilabel classification using IndoBERT embedding and MBERT classification. International Journal of Electrical and Computer Engineering, 14(1), 1071–1078. https://doi.org/10.11591/ijece.v14i1.pp1071-1078

Olthof, A. W., van Ooijen, P. M. A., & Cornelissen, L. J. (2021). Deep Learning-Based Natural Language Processing in Radiology: The Impact of Report Complexity, Disease Prevalence, Dataset Size, and Algorithm Type on Model Performance. Journal of Medical Systems, 45(10), 1–16. https://doi.org/10.1007/s10916-021-01761-4

O’Shaughnessy, D. (2026). An Overview of Recent Advances in Natural Language Processing for Information Systems. Applied Sciences, 16(2), 1122. https://doi.org/10.3390/app16021122

Patel, P. (2025). Infrastructure Economics of Sparse Mixture-of-Experts in Cloud-Native NLP: Benchmarking Cost, Accuracy, and Performance. Global Business & Economics Journal, 27(1), 1–18. https://doi.org/10.70924/f83n6wqz/vboj68ls

Pogiatzis, A., & Samakovitis, G. (2020). Using bilstm networks for context-aware deep sensitivity labelling on conversational data. Applied Sciences (Switzerland), 10(24), 1–17. https://doi.org/10.3390/app10248924

Qiu, Y., Dong, L., Zhang, W., Xing, H., & Huang, J. (2025a). A diffusion enhanced CRF and BiLSTM framework for accurate entity recognition. Scientific Reports, 15(1), 1–25. https://doi.org/10.1038/s41598-025-04036-x

Qiu, Y., Dong, L., Zhang, W., Xing, H., & Huang, J. (2025b). A diffusion enhanced CRF and BiLSTM framework for accurate entity recognition. Scientific Reports, 15(1). https://doi.org/10.1038/s41598-025-04036-x

Rahim, A., Zhong, Y., Ahmad, T., Ahmad, S., Pławiak, P., & Hammad, M. (2023). Enhancing Smart Home Security: Anomaly Detection and Face Recognition in Smart Home IoT Devices Using Logit-Boosted CNN Models. Sensors, 23(15), 1–42. https://doi.org/10.3390/s23156979

Salmani, M., Ghafouri, S., Sanaee, A., Razavi, K., Mühlhäuser, M., Doyle, J., Jamshidi, P., & Sharifi, M. (2023). Reconciling High Accuracy, Cost-Efficiency, and Low Latency of Inference Serving Systems. EuroMLSys ’23: Proceedings of the 3rd Workshop on Machine Learning and Systems, 78–86. https://doi.org/10.1145/3578356.3592578

Shidik, G. F., Saputra, F. O., Saraswati, G. W., Winarsih, N. A. S., Rohman, M. S., Pramunendar, R. A., Kusuma, E. J., Ratmana, D. O., Venus, V., Andono, P. N., & Hasibuan, Z. A. (2024). Indonesian disaster named entity recognition from multi source information using bidirectional LSTM (BiLSTM). Journal of Open Innovation: Technology, Market, and Complexity, 10(3), 1–12. https://doi.org/10.1016/j.joitmc.2024.100358

Singgalen, Y. A. (2025). Performance Analysis of IndoBERT for Sentiment Classification in Indonesian Hotel Review Data. Journal of Information System Research (JOSH), 6(2), 976–986. https://doi.org/10.47065/josh.v6i2.6505

Surya Suwardi Ansyah, A., Oranova Siahaan, D., Rizqi Paradisiaca Darnoto, B., Nopember, S., Studi Informatika, P., Ilmu Komputer, F., Jember, U., Krajan Timur, J., Sumbersari, K., Jember, K., & Timur, J. (2025). Integrated Named Entity Recognition and Identical-Entity Detection for Extracting Unique Information Sources in News Articles. Jurnal Teknologi Informasi Dan Komunikasi, 16(2), 72–83. https://doi.org/10.31849/digitalzone.v16i2

Wang, J., Yue, K., & Duan, L. (2023). Models and Techniques for Domain Relation Extraction: A Survey. In Journal of Data Science and Intelligent Systems (Vol. 1, Number 2, pp. 65–82). Bon View Publishing Pte Ltd. https://doi.org/10.47852/bonviewJDSIS3202973

Wang, S., Sun, X., Li, X., Ouyang, R., Wu, F., Zhang, T., Li, J., & Wang, G. (2025). Findings of the Association for Computational GPT-NER: Named Entity Recognition via Large Language Models. https://github.com/

Zhang, Y., & Xiao, G. (2024). Named Entity Recognition Datasets: A Classification Framework. In International Journal of Computational Intelligence Systems (Vol. 17, Number 1). Springer Science and Business Media B.V. https://doi.org/10.1007/s44196-024-00456-1

Zhu, Y. (2024). A knowledge graph and BiLSTM-CRF-enabled intelligent adaptive learning model and its potential application. Alexandria Engineering Journal, 91(1), 305–320. https://doi.org/10.1016/j.aej.2024.02.011




DOI: https://doi.org/10.31764/jtam.v10i3.38241

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Nurul Isnaeni Rahmat, Emha Taufiq Luthfi

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

_______________________________________________

JTAM already indexing:

                     


_______________________________________________

 

Creative Commons License

JTAM (Jurnal Teori dan Aplikasi Matematika) 
is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License

______________________________________________

_______________________________________________

_______________________________________________ 

JTAM (Jurnal Teori dan Aplikasi Matematika) Editorial Office: