Leveraging Test-Driven Prompt Engineering for Program Code Generation

Billdan Satriana Roseandree, Dwi Gumarang Shakti, Nadia Aqmarina Ghaisany, Ivan Jaelani Besti, Indira Syawanodya

Abstract


The rapid development of artificial intelligence and machine learning in today’s era has become inevitable, and humanity ought to embrace it. Generative AI is an important peak that disrupts the research trend in the information technology field. ChatGPT and Meta LLaMa are among the state-of-the-art generative AIs that has been widely used in various fields, including content generation and even code generation. Previous research on program code generation by generative AI has found that source codes that generated by LLMs are still inaccurate. This research aims to improve the correctness of the generated source code by implementing the test-driven development concept to enhance prompt engineering for program code generation – test-driven prompt engineering is the approach for the innovation. This study experiments two state-of-the-art generative AI by comparing their code quality that they produce. The results shows that both LLMs have similar code complexity and this is indicated by how both LLMs have the same number of lines of code and cyclomatic complexity, but differs in the maintainability index, where ChatGPT prevails over Meta LLaMa with maintainability index of 162 obtained by the code generated by ChatGPT, meanwhile Meta LLaMa achieved a lower MI of 145.


Keywords


Generative AI; Test-driven development; Prompt engineering; ChatGPT; Meta LLaMa;

References


Agha, D., Sohail, R., Meghji, A. F., Qaboolio, R., & Bhatti, S. (2023). Test Driven Development and Its Impact on Program Design and Software Quality: A Systematic Literature Review. VAWKUM Transactions on Computer Sciences, 11(1), 268–280. https://doi.org/10.21015/VTCS.V11I1.1494

Barenkamp, M., Rebstadt, J., & Thomas, O. (2020). Applications of AI in classical software engineering. AI Perspectives 2020 2:1, 2(1), 1–15. https://doi.org/10.1186/S42467-020-00005-4

Beck, K. (2002). Test Driven Development: By Example. In AddisonWesley Longman.

Feuerriegel, S., Hartmann, J., Janiesch, C., & Zschech, P. (2024). Generative AI. Business and Information Systems Engineering, 66(1), 111–126. https://doi.org/10.1007/S12599-023-00834-7/TABLES/2

Fui-Hoon Nah, F., Zheng, R., Cai, J., Siau, K., & Chen, L. (2023). Generative AI and ChatGPT: Applications, challenges, and AI-human collaboration. Journal of Information Technology Case and Application Research, 25(3), 277–304. https://doi.org/10.1080/15228053.2023.2233814

Booth, P. (n.d.). GitHub - escomplex/complexity-report: **UNMAINTAINED** Software complexity analysis for JavaScript projects. Retrieved June 5, 2024, from https://github.com/escomplex/complexity-report

Usta, S. (2020). GitHub - selcukusta/codalyze-rest-api: Codalyze: Code Complexity REST API. https://github.com/selcukusta/codalyze-rest-api

Hello GPT-4o | OpenAI. (2024, May 13). https://openai.com/index/hello-gpt-4o/

Heričko, T., & Šumak, B. (2023). Exploring Maintainability Index Variants for Software Maintainability Measurement in Object-Oriented Systems. Applied Sciences 2023, Vol. 13, Page 2972, 13(5), 2972. https://doi.org/10.3390/APP13052972

Idrisov, B., & Schlippe, T. (2024). Program Code Generation with Generative AIs. Algorithms 2024, Vol. 17, Page 62, 17(2), 62. https://doi.org/10.3390/A17020062

Introducing Meta Llama 3: The most capable openly available LLM to date. (2024, April 18). https://ai.meta.com/blog/meta-llama-3/

Jest · 🃏 Delightful JavaScript Testing. (n.d.). Retrieved June 5, 2024, from https://jestjs.io/

Khurana, D., Koli, A., Khatter, K., & Singh, S. (2023). Natural language processing: state of the art, current trends and challenges. Multimedia Tools and Applications, 82(3), 3713–3744. https://doi.org/10.1007/S11042-022-13428-4/FIGURES/3

Mccabe, T. J. (1976). A Complexity Measure. IEEE Transactions on Software Engineering, SE-2(4), 308–320. https://doi.org/10.1109/TSE.1976.233837

Pang, S., Nol, E., & Heng, K. (2024). ChatGPT-4o for English language teaching and learning: Features, applications, and future prospects. SSRN Electronic Journal. https://doi.org/10.2139/SSRN.4837988

Papis, B., Grochowski, K., Subzda, K., & Sijko, K. (2022). Experimental Evaluation of Test-Driven Development With Interns Working on a Real Industrial Project. IEEE Transactions on Software Engineering, 48(5). https://doi.org/10.1109/TSE.2020.3027522

Roman, A., & Mnich, M. (2021). Test-driven development with mutation testing – an experimental study. Software Quality Journal, 29(1). https://doi.org/10.1007/s11219-020-09534-x

Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Ellen Tan, X., Adi, Y., Liu, J., Sauvestre, R., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Canton Ferrer, C., Grattafiori, A., Xiong, W., Défossez, A., … Synnaeve, G. (2023). Code Llama: Open Foundation Models for Code. https://github.com/facebookresearch/codellama

Sonko, S., Adewusi, A. O., Obi, O. C., Onwusinkwue, S., Atadoga, A., Sonko, S., Adewusi, A. O., Obi, O. C., Onwusinkwue, S., & Atadoga, A. (2024). A critical review towards artificial general intelligence: Challenges, ethical considerations, and the path forward. Https://Wjarr.Com/Sites/Default/Files/WJARR-2024-0817.Pdf, 21(3), 1262–1268. https://doi.org/10.30574/WJARR.2024.21.3.0817

Stivers, T., Enfield, N. J., Brown, P., Englert, C., Hayashi, M., Heinemann, T., Hoymann, G., Rossano, F., De Ruiter, J. P., Yoon, K. E., & Levinson, S. C. (2009). Universals and cultural variation in turn-taking in conversation. Proceedings of the National Academy of Sciences of the United States of America, 106(26), 10587–10592. https://doi.org/10.1073/PNAS.0903616106

Stokel-Walker, C., & Van Noorden, R. (2023). What ChatGPT and generative AI mean for science. Nature, 614(7947), 214–216. https://doi.org/10.1038/D41586-023-00340-6

Sun, S., Zhang, Y., Yan, J., Gao, Y., Ong, D., Chen, B., & Su, J. (2023). Battle of the Large Language Models: Dolly vs LLaMA vs Vicuna vs Guanaco vs Bard vs ChatGPT - A Text-to-SQL Parsing Comparison. Findings of the Association for Computational Linguistics: EMNLP 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.750

Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., & Lample, G. (2023). LLaMA: Open and Efficient Foundation Language Models. https://arxiv.org/abs/2302.13971v1

Wang, J., Shi, E., Yu, S., Wu, Z., Ma, C., Dai, H., Yang, Q., Kang, Y., Wu, J., Hu, H., Yue, C., Zhang, H., Liu, Y., Pan, Y., Liu, Z., Sun, L., Li, X., Ge, B., Jiang, X., … Zhang, S. (2023). Prompt Engineering for Healthcare: Methodologies and Applications. https://arxiv.org/abs/2304.14670v2

Zhang, H. (2009). An Investigation of the Relationships between Lines of Code and Defects. https://doi.org/10.1109/ICSM.2009.5306304

Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., & Ba, J. (2022). Large Language Models Are Human-Level Prompt Engineers. https://arxiv.org/abs/2211.01910v2




DOI: https://doi.org/10.17509/seict.v6i2.70731

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Journal of Software Engineering, Information and Communication Technology (SEICT)

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

Journal of Software Engineering, Information and Communicaton Technology (SEICT), 
(e-ISSN:
2774-1699 | p-ISSN:2744-1656) published by Program Studi Rekayasa Perangkat Lunak, Kampus UPI di Cibiru.


 Indexed by.