Automated multiple-choice question generation for soft skills development using LLMs: a short narrative review
Abstract
The development of soft skills is crucial in higher education and professional training, yet traditional assessment methods often struggle to provide personalized and scalable solutions. This narrative review explores the potential of Large Language Models (LLMs), specifically GPT-based models, in generating multiple-choice questions (MCQs) aimed at soft skills development. A structured analysis of recent adjacent literature suggests that LLMs may offer considerable promise for automating the creation of diverse, contextually relevant, and construct-specific MCQs across various domains, including psychological assessments, medical training, and educational settings. Key techniques such as prompt engineering, model fine-tuning, and retrieval-augmented generation are identified as effective methods for enhancing the accuracy and relevance of AI-generated questions. However, challenges such as biases, factual inaccuracies, and the need for human oversight remain significant concerns. The reviewed studies consistently highlight the importance of integrating human expertise to ensure the validity, fairness, and quality of the generated content. Despite these challenges, the use of LLMs could streamline the creation of personalized learning tools, offering scalable solutions for soft skills assessment. Future research should focus on refining AI techniques, developing robust evaluation frameworks, and exploring the impact of AI-generated MCQs on learners' skill development. This review emphasizes the need for a hybrid approach, combining AI's efficiency with expert validation, to fully realize the potential of LLMs in soft skills education.
References
- G. W. Mitchell, L. B. Skinner, and B. J. White, “Essential soft skills for success in the twenty-first century workforce as perceived by business educators,” Delta Pi Epsilon Journal, vol. 52, no. 1, 2010.
- A. F. Hendarman and U. Cantner, “Soft skills, hard skills, and individual innovativeness,” Eurasian Business Review, vol. 8, pp. 139–169, 2018. doi: 10.1007/s40821-017-0076-6.
- B. Schulz, “The importance of soft skills: Education beyond academic knowledge,” Journal of Language and Communication, 2008.
- L. G. Malicay, “The integration of soft skills in professional education: Exploring the importance of communication, teamwork, and interpersonal skills in professional training and the methods used to incorporate them into educational programs,” International Journal of Advanced Research in Science, Communication and Technology, vol. 3, no. 1, pp. 836–843, 2023. doi: 10.48175/IJARSCT-11967.
- S. Maity and A. Deroy, “The future of learning in the age of generative AI: Automated question generation and assessment with large language models,” arXiv preprint arXiv:2410.09576, 2024. doi: 10.48550/arXiv.2410.09576.
- F. S. Mohammed and F. Ozdamli, “A systematic literature review of soft skills in information technology education,” Behavioral Sciences, vol. 14, no. 10, p. 894, 2024. doi: 10.3390/bs14100894.
- G. Fontaine, S. Cossette, M. A. Maheu-Cadotte, T. Mailhot, M. F. Deschênes, and G. Mathieu-Dupuis, “Effectiveness of adaptive e-learning environments on knowledge, competence, and behavior in health professionals and students: Protocol for a systematic review and meta-analysis,” JMIR Research Protocols, vol. 6, no. 7, pp. 1–10, 2017. doi: 10.2196/resprot.8085.
- M. J. Page, J. E. McKenzie, P. M. Bossuyt, I. Boutron, T. C. Hoffmann, C. D. Mulrow, L. Shamseer, J. M. Tetzlaff, E. A. Akl, S. E. Brennan et al., “The PRISMA 2020 statement: An updated guideline for reporting systematic reviews,” BMJ, vol. 372, 2021. doi: 10.1136/bmj.n71.
- R. Ferrari, “Writing narrative style literature reviews,” Medical Writing, vol. 24, no. 4, pp. 230–235, 2015. doi: 10.1179/2047480615Z.000000000329.
- D. Torres-Salinas, R. Ruiz-Pérez, and E. Delgado-López-Cózar, “Google Scholar como herramienta para la evaluación científica,” Profesional de la Información, vol. 18, no. 5, pp. 501–510, 2009. doi: 10.3145/EPI.2009.SEP.03.
- G. Halevi, H. Moed, and J. Bar-Ilan, “Suitability of Google Scholar as a source of scientific information and as a source of data for scientific evaluation—review of the literature,” Journal of Informetrics, vol. 11, no. 3, pp. 823–834, 2017. doi: 10.1016/j.joi.2017.06.005.
- D. Moher, A. Liberati, J. Tetzlaff, and D. G. Altman, “Preferred reporting items for systematic reviews and meta-analyses: The PRISMA statement,” PLoS Medicine, vol. 6, no. 7, p. e1000097, 2009. doi: 10.1371/journal.pmed.1000097.
- J. P. Higgins and S. Green, Cochrane Handbook for Systematic Reviews of Interventions. John Wiley & Sons, 2011. doi: 10.1002/9780470712184.
- A. Dobrescu, B. Nussbaumer-Streit, I. Klerings, G. Wagner, E. Persad, I. Sommer, H. Herkner, and G. Gartlehner, “Restricting evidence syntheses of interventions to English-language publications is a viable methodological shortcut for most medical topics: A systematic review,” Journal of Clinical Epidemiology, vol. 137, pp. 209–217, 2021. doi: 10.1016/j.jclinepi.2021.04.012. Available: https://linkinghub.elsevier.com/retrieve/pii/S0895435621001347
- D. Gough, S. Oliver, and J. Thomas, An Introduction to Systematic Reviews. SAGE, 2017. doi: 10.4135/9781036234942.
- M. Callaham, R. L. Wears, and E. Weber, “Journal prestige, publication bias, and other characteristics associated with citation of published studies in peer-reviewed journals,” JAMA, vol. 287, no. 21, pp. 2847–2850, 2002. doi: 10.1001/jama.287.21.2847.
- E. C. Martinez, P. E. G. Hasbun, V. P. S. Vargas, O. Y. García-González, M. D. F. Madera, D. E. R. Capistrán, T. C. Carmona, C. S. Cruz, and C. T. Hooper, “A comprehensive guide to conduct a systematic review and meta-analysis in medical research,” Medicine, vol. 104, p. e41868, 2025. doi: 10.1097/MD.0000000000041868. Available: https://journals.lww.com/10.1097/MD.0000000000041868
- N. R. Haddaway, P. Woodcock, B. Macura, and A. Collins, “Making literature reviews more reliable through application of lessons from systematic reviews,” Conservation Biology, vol. 29, no. 6, pp. 1596–1605, 2015. doi: 10.1111/cobi.12541.
- A. V. Kulkarni, B. Aziz, I. Shams, and J. W. Busse, “Comparisons of citations in Web of Science, Scopus, and Google Scholar for articles published in general medical journals,” JAMA, vol. 302, no. 10, pp. 1092–1096, 2009. doi: 10.1001/jama.2009.1307.
- P. Lee, S. Fyffe, M. Son, Z. Jia, and Z. Yao, “A paradigm shift from ‘human writing’ to ‘machine generation’ in personality test development: An application of state-of-the-art natural language processing,” Journal of Business and Psychology, vol. 38, pp. 163–190, 2023. doi: 10.1007/s10869-022-09864-6. Available: https://doi.org/10.1007/s10869-022-09864-6
- F. M. Götz, R. Maertens, S. Loomba, and S. van der Linden, “Let the algorithm speak: How to use neural networks for automatic item generation in psychological scale development,” Psychological Methods, vol. 29, pp. 494–518, 2024. doi: 10.1037/met0000540.
- B. E. Hommel, F. J. M. Wollang, V. Kotova, H. Zacher, and S. C. Schmukle, “Transformer-based deep neural language modeling for construct-specific automatic item generation,” Psychometrika, vol. 87, pp. 749–772, 2022. doi: 10.1007/s11336-021-09823-9. Available: https://doi.org/10.1007/s11336-021-09823-9
- U. Lee, H. Jung, Y. Jeon, Y. Sohn, W. Hwang, J. Moon, and H. Kim, “Few-shot is enough: Exploring ChatGPT prompt engineering method for automatic question generation in English education,” Education and Information Technologies, vol. 29, pp. 11483–11515, 2024. doi: 10.1007/s10639-023-12249-8. Available: https://doi.org/10.1007/s10639-023-12249-8
- Y. S. Kıyak, Özlem Coşkun, İrem Budakoğlu, and C. Uluoğlu, “ChatGPT for generating multiple-choice questions: Evidence on the use of artificial intelligence in automatic item generation for a rational pharmacotherapy exam,” European Journal of Clinical Pharmacology, vol. 80, pp. 729–735, 2024. doi: 10.1007/s00228-024-03649-x.
- I. R. Indran, P. Paranthaman, N. Gupta, and N. Mustafa, “Twelve tips to leverage AI for efficient and effective medical question generation: A guide for educators using ChatGPT,” Medical Teacher, vol. 46, pp. 1021–1026, 2024. doi: 10.1080/0142159X.2023.2294703.
- Y. Attali, A. Runge, G. T. Laflair, K. Yancey, S. Goodwin, Y. Park, and A. A. V. Davier, “The interactive reading task: Transformer-based automatic item generation,” Frontiers in Artificial Intelligence, vol. 5, 2022. doi: 10.3389/frai.2022.903077.
- K. Soman, P. W. Rose, J. H. Morris, R. E. Akbas, B. Smith, B. Peetoom, C. Villouta-Reyes, G. Cerono, Y. Shi, A. Rizk-Jackson, S. Israni, C. A. Nelson, S. Huang, and S. E. Baranzini, “Biomedical knowledge graph-optimized prompt generation for large language models,” Bioinformatics, vol. 40, 2024. doi: 10.1093/bioinformatics/btae560. Available: http://arxiv.org/abs/2311.17330
- R. Circi, J. Hicks, and E. Sikali, “Automatic item generation: Foundations and machine learning-based approaches for assessments,” Frontiers in Education, 2023. doi: 10.3389/feduc.2023.858273.
- D. Pugh, A. D. Champlain, M. Gierl, H. Lai, and C. Touchie, “Can automated item generation be used to develop high quality MCQs that assess application of knowledge?” Research and Practice in Technology Enhanced Learning, vol. 15, 2020. doi: 10.1186/s41039-020-00134-8.
License
This article is licensed under Creative Commons Attribution 4.0 International License (CC BY 4.0)