A Statistical Machine Translation Approach for Banjarese–Indonesian Language Translation
Abstract
The Banjarese language has recently become a particular concern not only among regional language observers in Kalimantan but also at the national and even international levels. This is due to the fact that regional languages are increasingly being abandoned by the younger generation as a result of the hegemony of contemporary languages and other languages that are perceived as more prestigious. This phenomenon can be clearly observed in the significant decline in the number of regional language speakers. The purpose of this study is to preserve the regional language through the development of a Banjarese–Indonesian and Indonesian–Banjarese machine translation system. This research is expected to serve as a reference for other linguistic researchers who face limitations in using the Banjarese language. Automatic machine translation has been widely applied in various fields such as language recognition and vocabulary learning, sentence tree analysis, question-answering systems, text summarization, sentiment analysis, and computational linguistics. This study employs the Statistical Machine Translation (SMT) method to translate each word in a sentence into Banjarese and vice versa. The proposed method is evaluated using the Statistical Calculation, METEOR (Metric for Evaluation of Translation with Explicit Ordering) Universal Tools and human justification. The test results using the statistical calculation show that the word translation success rate using the Statistical Machine Translation method reaches 48%. Meanwhile, the evaluation of the naturalness of translated sentences using the METEOR Universal Tools with 1,000 sentence data shows that the proposed method achieves a final_score of 80. The final evaluation was conducted using human judgment involving 15 language experts. The results indicate that the proposed method is able to translate sentences based on individual words and the highest word probability available without altering words based on synonym lists. The value of naturalness evaluation using human judgment are 72.6%.
Keywords
References
Alvarez-Carmona, M., Aranda, R., Rodríguez-Gonzalez, A. Y., Fajardo-Delgado, D., Sánchez, M. G., Pérez-Espinosa, H., et al. (2022). Natural language processing applied to tourism research: A systematic review and future research directions. Journal of King Saud University - Computer and Information Sciences, 34(10), 10125–10144.
Aqlan, A. A. Q., Manjula, B., & Lakshman Naik, R. (2019). A study of sentiment analysis: Concepts, techniques, and challenges. In Lecture Notes on Data Engineering and Communications Technologies (Vol. 28, pp. 147–162). Springer. http://dx.doi.org/10.1007/978-981-13-6459-4_16
Barmawi, A. M., & Muhammad, A. (2019). Paraphrasing method based on contextual synonym substitution. Jurnal ICT Research and Applications, 13(3), 257–282.
Barmawi, A. M., Wahyudi, B. A., & Pristi, T. (2023). Linguistic based one time password. International Journal of Electrical Engineering and Informatics, 15(1), 1–16.
Fashwan, A., & Alansary, S. (2021). A morphologically annotated corpus and a morphological analyzer for Egyptian Arabic. Procedia CIRP, 189, 203–210. https://doi.org/10.1016/j.procs.2021.05.084
Ghirrid, A. A., Sari, R. T. K., & Aldisa, R. T. (2024). Algoritma natural language processing untuk aplikasi penerjemah (Indonesia–Jawa) menggunakan metode speech processing. Jurnal JTIK (Jurnal Teknologi Informasi dan Komunikasi), 8(3), 746–759. https://doi.org/10.35870/jtik.v8i3.2244
Guinovart, X. G. (2019). Enriching parallel corpora with multimedia and lexical semantics from the CLUVI corpus to WordNet and SemCor. In John Benjamins Publishing Company (pp. 141–158).
Hapsari, W. P., Labib, U. A., Haryanto, H., & Safitri, D. W. (2021). A literature review of Human, Organization, Technology (HOT) – Fit evaluation model. In Proceedings of the 6th International Seminar on Science Education (ISSE 2020) (pp. 876–883).
Hasmianti, L., Usman, U., & Amir, J. (2023). Pergeseran penggunaan kata sapaan oleh generasi milenial Banjar di Kota Banjarmasin. JP-BSI: Jurnal Pendidikan Bahasa dan Sastra Indonesia, 8(2), 122.
Hermawan, D. (2022). Pergeseran penggunaan bahasa. Locana, 5(1), 23–37.
Kamariah, & Jamiatul, N. K. H. (2023). Konservasi Bahasa Banjar sebagai usaha. Konfiks, 10(2), 24. https://journal.unismuh.ac.id/index.php/konfiksPermalink/DOI:https://doi.org/10.26618/jk/13118
Larasati, S. D. (2012). IDENTIC corpus: Morphologically enriched Indonesian-English parallel corpus. In Proceedings of the 8th International Conference on Language Resources and Evaluation (pp. 902–906).
Liu, B., & Huang, L. (2021). ParaMed: A parallel corpus for English–Chinese translation in the biomedical domain. BMC Medical Informatics and Decision Making, 21(1).
Lopez, A. (2023). Machine translation evaluation metrics benchmarking: From traditional MT to LLMs (1st ed.). Universitat de Barcelona.
Mohammed, T. A. S. (2022). The use of corpora in translation into the second language: A project-based approach. Frontiers in Education, 7, 1–14.
Muhammad, A., & Kamariah, K. (2020). Pengurai kalimat Bahasa Banjar dengan menggunakan parser PC-PATR. Jurnal Linguistik Komputasional, 3(1), 20.
Muhammad, A., & Widyastuti, N. (2024). Pengembangan aplikasi part-of-speech tagger Bahasa Banjar menggunakan metode pengembangan DevOps. JIKOMTI: Jurnal Ilmiah Ilmu Komputer dan Teknologi Informasi, 1(1).
Muhammad, A., Winda, N., Firizkiansah, A., Setiawan, D., Dewi, S. H. F., Maulana, I. R., Ardiansyah, M. (2025). Review of Banjarnese Neural Machine Translation Development With Minimal Resources. Journal of Software Engineering, Information and Communicaton Technology (SEICT), Vol 6, No 1.
Muttaqin, A. I. (2019). Konstruksi verba gerak direksional dalam Bahasa Banjar. PRASASTI: Journal of Linguistics, 4(2), 99–103. https://jurnal.uns.ac.id/pjl/article/view/34129
Nur, S., Assyifa, A. N., & Nurjannah, H. (2023). Pengembangan aplikasi penerjemah bahasa isyarat Indonesia (Bisindo) menggunakan metode long-short term memory. EDUSAINTEK: Jurnal Pendidikan, Sains dan Teknologi, 11(1), 13–30.
Oliver, A., & Álvarez, S. (2024). LitPC: A set of tools for building parallel corpora from literary works. In CTT 2024 – 1st Workshop on Creative Translation Technologies Proceedings (pp. 21–31).
Pan, B., & Qin, Q. (2022). Construction of parallel corpus for English translation teaching based on computer aided translation software. Computer-Aided Design and Applications, 19(s1), 70–80.
Prabowo, A., & Indra Sanjaya, F. (2024). Penerapan metode transfer learning pada IndoBERT untuk analisis sentimen teks Bahasa Jawa Ngoko Lugu. Simkom, 9(2), 205–217.
Purnajaya, A., Indriani, F., & Faisal, M. R. (2021). Pengenalan suara pada kamus Banjar–Indonesia dan Indonesia–Banjar menggunakan statistik inferensi. Jurnal Ilmiah Informatika, 8(1), 1–8. https://doi.org/10.33884/jif.v8i01.1727
Pranaja, M. A., & Nurhidayat, A. I. (2023). Pembuatan aplikasi mesin penerjemah menggunakan metode No Language Left Behind dari bahasa Indonesia ke bahasa Banjar. Jurnal Manajemen Informatika, 12(1).
Rui, L., & Xiuli, G. (2022). Basic research on construction of multimodal parallel corpus of tourism translation in new media era. Academic Journal of Humanities and Social Sciences, 5(15), 139–144.
Satanjeev Banerjee and Alon Lavie. (2005). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the Workshop on Intrinsic and Extrinsic Evaluation Measures for MT and/or Summarization (EvalMT '05), Association for Computational Linguistics, Stroudsburg, PA, USA, 65–72.
Setiawan, D., Muhammad, A., & Firizkiansah, A. (2024). Pengklasifikasian dokumen teks Bahasa Indonesia berbasis vector space model dengan menggunakan metode k-nearest neighbor (k-NN) dan Euclidean distance. JIKOMTI: Jurnal Ilmiah Ilmu Komputer dan Teknologi Informasi, 1(1).
Shen, N. (2022). English–Chinese corpus collection and translation wisdom algorithm implementation based on Ajax + JQuery. International Journal of Science and Engineering Applications, 11(12), 300–302.
Spatioti, A. G., Kazanidis, I., & Pange, J. (2022). A comparative study of the ADDIE instructional design model in distance education. Information, 13(9), 1–20.
Susilawati, E., Winda, N., & Akbari, S. (2021). Representatif model pemimpin masyarakat Banjar pada cerita rakyat Kisah Datu Wani. In Prosiding SENSASEDA (Vol. 1, pp. 55–64). https://jurnal.stkipbjm.ac.id/index.php/-sensaseda/article/download/1560/786
Team, Si Palui., Si Palui, Banjarmasin Post. https://banjarmasin.-tribunnews.com/, 2020-2025.
Team, Harian Kompas, X. https://x.com/hariankompas. 2020-2025.
Valentino, M., Thayaparan, M., & Freitas, A. (2021). Unification-based reconstruction of multi-hop explanations for science questions. In EACL 2021 – Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics (pp. 200–211).
Wahyudinata, T., Sujaini, H., & Nyoto, R. D. (2016). Implementasi mesin penerjemah statistik berbasis Android dengan Moses Decoder. Jurnal Sistem dan Teknologi Informasi (JUSTIN), 3(1).
Winda, N., & Muhammad, A. (2023). Pengembangan parsing PCPATR sebagai preservasi Bahasa dan Sastra Banjar. Bahasa dan Sastra, 9. https://e-journal.my.id/onoma
Zakiah, Sumarlam, & Supana. (2022). Faktor pemertahanan Bahasa Banjar pada pedagang pasar terapung di Kota Banjarmasin. In Proceedings Seminar Nasional Linguistik dan Sastra (Vol. 4, pp. 1–9). https://jurnal.uns.ac.id/-prosidingsemantiks
Zheng, X., & Wu, H. (2022). Autoregressive linguistic steganography based on BERT and consistency coding. Security and Communication Networks, 2022, 1–8.
DOI: https://doi.org/10.17509/seict.v6i2.93853
Refbacks
- There are currently no refbacks.
Copyright (c) 2026 Journal of Software Engineering, Information and Communication Technology (SEICT)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Journal of Software Engineering, Information and Communicaton Technology (SEICT),
(e-ISSN:2774-1699 | p-ISSN:2744-1656) published by Program Studi Rekayasa Perangkat Lunak, Kampus UPI di Cibiru.
Indexed by.



