Artificial intelligence (AI) has recently been integrated into language teaching and assessment, revolutionizing the evaluation of language skills. Therefore, this study aims to explore the perceptions of university English language teachers towards AI-generated grammar assessments. It identifies the differences between AI-generated and human-made tests, reveals the potential benefits and limitations of AI-generated items and proposes practical strategies to overcome these limitations.
The data were collected from 12 experienced English as a Foreign Language (EFL) teachers at Saudi Electronic University, using a mixed-methods approach that comprised a test evaluation checklist and semi-structured interviews. The checklist, adapted from Brown (2004) and Brown (2005), evaluated the test items in terms of validity, reliability, contextuality, involvement and quality.
The results revealed that the human-made test excelled in validity, contextuality and involvement, while the AI-generated test demonstrated strengths in reliability and clarity. The thematic analysis of the interviews’ responses identified three benefit-related themes: time-saving, question variety and personalization; five challenge-related themes: ambiguity, poor prompts, over-reliance, ethical concerns and errors; and two suggested strategies: monitoring and human editing.
While using AI has promising potential, its effective implementation in language classrooms requires careful planning. Thus, in light of the findings, the study proposes practical strategies to bring out the strengths of AI while mitigating its limitations.
Introduction
Research on the impact of AI tools in education is still in its early stages. However, its significance cannot be underestimated, as it has already improved teaching and learning through efficiency and curriculum personalization (Hwang and Chang, 2023). Pokrivcakova (2019) noted that AI integration in education elevates both learning and teaching by creating a cutting-edge learning environment where teaching can be more adaptable and learning can be more personalized. Yet, AI presents several threats that need to be addressed, including concerns regarding ethics, assessment, privacy, security, disruption of the teacher-student relationship, harm to students' character development, and disruption of social interactions.
For second language learners, Jia et al. (2022) found that AI offers an authentic and ubiquitous learning environment. This technological integration facilitates the entire learning experience, making it more engaging and economical. In the context of language assessment, AI tools, such as automated scoring systems, intelligent tutoring platforms, and generative language models, have shown promise in enhancing efficiency, providing immediate feedback, and supporting personalized learning pathways (Stošić and Malyuga, 2024). However, most current research tends to investigate AI’s role in assessing writing and speaking skills, with less attention given to grammar assessment (Crompton et al., 2024; Lalira et al., 2024). This lack of empirical investigation creates uncertainty about the effectiveness, reliability, and contextual appropriateness of AI-generated grammar assessments, particularly in university-level language classrooms (O, 2024). Therefore, this study explores how English language instructors perceive and evaluate AI-generated grammar tasks, as well as how these tools can be optimally utilized or adapted (Maeda, 2024).
Additionally, this study draws on the well-established principles of language assessment proposed by Brown (2004, 2005), whose frameworks emphasize core dimensions: validity, reliability, contextuality, involvement, and quality. These principles are widely recognized in second language assessment as essential criteria for evaluating the test quality. They serve as a theoretical lens through which the strengths and limitations of AI-generated and human-made grammar tests are interpreted, reinforcing the study’s aim to critically compare AI- and human-generated assessments against well-established language testing standards.
Literature review
Advantages and challenges of AI in language assessment
In recent years, AI chatbots have become increasingly prominent in language teaching and assessment. Tools such as ChatGPT, DeepSeek, Jasper, and Google Bard are designed to process natural language and generate coherent, human-like responses. While they offer promising benefits in language testing, such as enhancing efficiency, consistency, and scalability, researchers caution against relying entirely on these tools without human moderation. For instance, AI-powered assessment tools can quickly generate diverse item types (e.g. multiple-choice, true/false, cloze tests) and adapt to learners’ levels, making them attractive for teachers with limited time (Alkhateeb et al., 2025). However, there are still concerns regarding validity, fairness, and contextual accuracy. AI tools may lack sensitivity to cultural and linguistic nuances, potentially resulting in test items that misalign with the course objectives or learner realities (Bulut et al., 2024). Although they have demonstrated impressive language generation capabilities, they are not yet suitable replacements for expert human judgment in assessment design, especially when the content of the test must reflect deep pedagogical goals. While AI has the potential to support language teachers in creating and delivering assessments, its current limitations emphasize the importance of hybrid models that combine AI with human expertise (Alkhateeb et al., 2025; Owan et al., 2023).
Potentials and pitfalls of integrating AI in grammar testing
Traditional grammar testing faces several key challenges relating to validity, reliability, test engagement, and relevance. Therefore, AI-powered tools offer promising solutions to address some of these issues. For example, applications in grammar testing have been found to produce rapid and consistent feedback, potentially enhancing both accuracy and test efficiency (Katsarou et al., 2025). Crompton et al. (2024) presented a comprehensive systematic review of the use of AI in English language teaching and assessment. The review identified the benefits of AI in teaching and assessing language skills, including grammar. It also highlighted several challenges, such as technology failures, limited capabilities, and the standardization of language. Lalira et al. (2024) investigated the effect of AI-assisted learning tools on grammar proficiency among students from various academic programs at two universities in Indonesia. The study concluded that AI tools, such as Grammarly and ChatGPT, significantly improve students' grammar proficiency. The authors recommended cautious supervision to avoid over-dependence and unethical use.
Additionally, Nindya and Taufiqulloh (2024) explored EFL learners' perceptions of grammar accuracy and the effect of AI-driven feedback tools on their self-assessment and academic achievement. The participants reported that AI tools identified missing errors and enhanced their engagement and grammar proficiency; however, a few still preferred the traditional feedback. The researchers recommended a blended method that combines AI and human interaction. Maeda (2024) investigated the use of AI models to field-test multiple-choice English grammar questions. The study assessed AI performance alongside human test-takers, identifying which items might be too easy, too hard, or potentially flawed. The findings revealed that AI can be an effective tool for pre-testing and generating grammar test items, enhancing efficiency and item quality in language assessments while reducing overall costs. Despite these advantages, AI may still struggle with handling nuanced grammar. Concerns have also been raised about fairness, transparency, and ethical use. Although AI can support adaptive test design and provide personalized feedback, its underlying mechanisms are often difficult to interpret, as they operate through opaque, proprietary systems (Bulut et al., 2024).
AI-generated vs human-made test items
There has been growing research on the role and effect of AI in language assessment, especially in providing feedback (Shi and Aryadoust, 2024). Yet, there is still a limited body of research that directly compares AI-generated tests with those created by human examiners (O, 2024). Several factors might contribute to this dearth of research. First, the integration of AI into language testing practices is relatively recent (Kic-Drgas and Kılıçkaya, 2024; O, 2024). Second, assessing language skills involves addressing complex and nuanced criteria that AI tools may not have fully grasped yet (Abida et al., 2023; Stоšić and Malyuga, 2024). However, given the rapid growth and wide acceptance of AI today, it is most likely that further research will yield better insights into the pros and cons of AI-generated tests in language assessment (Ratnayanti et al., 2023; Stоšić and Malyuga, 2024).
Research has found that human examiners bring years of knowledge, contextual understanding, and the ability to capture the nuances of language. Thus, they can create valid and contextually relevant assessments that align with the content, learning objectives, and students’ needs (Chan and Louisa, 2023; Owan et al., 2023). However, inconsistencies and ambiguities might be detected in human-made assessments due to several factors, such as human error, subjectivity, and biases that might affect the wording of questions.
On the other hand, AI was reported to have high reliability, consistency, language clarity, as well as the ability to tailor assessments to students’ proficiency (Stоšić and Malyuga, 2024). However, since AI tools may have limited understanding of context and the nuances of language that human assessors possess, a collaborative approach that combines AI and the experience of examiners can optimize assessment practices (Abida et al., 2023; Kic-Drgas and Kılıçkaya, 2024; Stоšić and Malyuga, 2024), and therefore “[It] can result in more accurate and comprehensive language evaluations” (Ratnayanti et al., 2023, p. 16).
Research problem and aim of the study
English language instructors face numerous challenges, including increasing workloads and time constraints, which affect not only their teaching practices but also their overall well-being and job satisfaction. One time-consuming responsibility is designing high-quality, level-appropriate grammar assessments, which demands both linguistic precision and contextual sensitivity. AI offers promising potential to ease instructors’ workload, particularly through the automated generation of grammar test items. AI-powered tools, such as ChatGPT, can rapidly create a variety of question types, which ease the burden of designing exercises, quizzes, and tests. However, research on the use of AI specifically for grammar instruction and assessment remains limited. Most existing studies have focused on AI applications in writing and speaking, leaving a gap in how AI can be utilized to support grammar assessment in language classrooms (Crompton et al., 2024; Lalira et al., 2024; O, 2024).
Therefore, this study aims to investigate the attitudes of university English language teachers towards AI-generated grammar assessments and examine their concerns regarding the quality, appropriateness, and practical application of such tools. Furthermore, the study seeks to identify both the potential benefits and limitations of AI-generated grammar tests and proposes practical strategies for addressing these challenges. By doing so, the study contributes to the emerging body of research on AI in language education, and it provides actionable insights for institutions and teacher-training programs aiming to incorporate AI tools more effectively and responsibly (Maeda, 2024).
Research questions
The study seeks to answer the following questions:
What are the differences between AI-generated and human-made tests?
What are the potential benefits of using AI-generated grammar exercises or tests in language classrooms?
What are the challenges reported by university English language teachers regarding using AI-generated grammar activities or tests?
What are some practical strategies to help language instructors tackle the limitations of AI-generated exercises?
Methodology
Research design
The study used a mixed-methods design to explore teachers' attitudes toward AI-generated test items. The research process began with the development and implementation of a test evaluation checklist, followed by semi-structured interviews. By integrating both quantitative and qualitative approaches, the study provided a more comprehensive understanding of the topic than a single-method approach (Creswell, 2012).
The study utilized two grammar exams: one created by human examiners and the other generated by AI. The human-generated exam consisted of multiple-choice questions, true/false statements, and open-ended questions, designed to test a wide range of grammatical concepts, including sentence types, verb tenses, modal auxiliary verbs, and error correction. The exam emphasized the identification of correct answers based on real-world, nuanced language understanding, providing clear and contextually relevant questions (see Supplementary Material 1). The AI-generated grammar test was developed using ChatGPT (OpenAI). To ensure alignment with the course objectives and fairness in comparison to the human-made test, the researchers used a carefully structured prompt that read: 'Generate a 10-item grammar quiz for intermediate university-level EFL students. Include multiple-choice, true/false, fill-in-the-blank, and short-answer formats. Focus on past simple vs past progressive, modal verbs (can, could, must, should), and dependent vs independent clauses. Ensure the questions are pedagogically appropriate for adult learners.' The output was reviewed for clarity, language accuracy, and alignment with learning objectives before being used for analysis (see Supplementary Material 2).
Participants
Data were collected from 12 experienced EFL instructors (nine females and three males) at Saudi Electronic University (SEU) in Saudi Arabia. They specialized in Linguistics and TESOL, with eight majoring in Linguistics and four in TESOL. Their academic ranks were distributed as follows: eight assistant professors, three lecturers, and one language instructor.
Instrumentation
The test evaluation checklist, designed by the researchers, was used to gather information on how participants would judge the quality of both AI-generated and human-made tests. Interviews were also conducted to gather the participants’ insights and perspectives on this topic after obtaining the ethical approval from the Ethics Committee at SEU (No. SEUREC-4606). The following section provides a detailed description of each instrument and the procedures in detail.
Test evaluation checklist
A comprehensive checklist was developed and used as a tool to evaluate AI-generated and human-made tests. It was carefully designed to assess several aspects of the tests and conduct a thorough evaluation. The checklist’s criteria and questions were adapted from Brown (2004, 2005) as their work provides a solid framework for evaluating language assessments (see Table 1). The newly designed checklist might also serve both future researchers and instructors. To facilitate replicable studies, it can be used as a standardized tool for assessing the quality of test items. Additionally, it can be utilized by instructors to critically evaluate their tests and ensure they meet high standards. The checklist contains five major criteria: validity, reliability, contextuality, involvement, and quality of items, each with detailed sub-criteria, and it uses the Likert scale with options of (Yes = 3), (Somewhat = 2), and (No = 1). Validity focuses on the items’ alignment with the learning objectives and their consistency with the course content. Reliability involves checking whether the items are suitable for the students’ language proficiency and cognitive levels; it assesses the clarity and conciseness of language. Contextuality examines the existence of keywords related to the applied grammar rules and assesses the contextualization of these rules. Involvement evaluates the topics’ relevance to students’ knowledge and the gradual presentation of the grammar rules. The Quality of Items is further divided into sub-categories of multiple-choice, True/False, fill-in, and short-answer items, and they are all analyzed for clarity, conciseness, and non-ambiguity.
The test evaluation checklist
| Criteria | Checklist questions | Yes 3 | Somewhat 2 | No 1 | |
|---|---|---|---|---|---|
| 1. Validity | 1. Are the items aligned with the learning objectives of the course? | ||||
| 2. Are the items consistent with what the students have learned in the course? | |||||
| 2. Reliability | 3. Are the questions suitable to the language proficiency and cognitive level of the students being assessed? | ||||
| 4. Is the language used to craft the questions clear, concise, and unambiguous? | |||||
| 3. Contextuality | 5. Does the grammar exercise contain keywords that show the type of the applied grammar rule? | ||||
| 6. Is the exercise contextualized enough to present the usage and the form of the grammar rule? | |||||
| 4. Involvement | 7. Does the exercise grab the interest of students through tackling topics relevant to their background? | ||||
| 8. Does the exercise present the grammar rule appropriately in a graded way (easy to difficult)? | |||||
| 5. Quality of Items | 5.1. Multiple-Choice Items | 9. Are the stems in the multiple-choice items clearly worded and free of ambiguity and irrelevant details? | |||
| 10. Are the alternatives/distractors in the multiple-choice items plausible and clear? | |||||
| 5.2. True/False Items | 11. Are the T/F items worded carefully enough so they can be judged without ambiguity? | ||||
| 12. Do the T/F items include only one main idea? | |||||
| 5.3. Fill-in Items | 13. Is there sufficient context in the fill-in items to convey the intent of the questions? | ||||
| 14. Are the required, correct responses in the fill-in items concise? | |||||
| 5.4. Short-Answer Items | 15. Are the short-answer items formatted so that only one relatively concise answer is possible? | ||||
| 16. Are the short-answer items written in clear and direct questions? | |||||
| Would you like to add any comments? | |||||
| Criteria | Checklist questions | Yes | Somewhat | No | |
|---|---|---|---|---|---|
| 1. Validity | 1. Are the items aligned with the learning objectives of the course? | ||||
| 2. Are the items consistent with what the students have learned in the course? | |||||
| 2. Reliability | 3. Are the questions suitable to the language proficiency and cognitive level of the students being assessed? | ||||
| 4. Is the language used to craft the questions clear, concise, and unambiguous? | |||||
| 3. Contextuality | 5. Does the grammar exercise contain keywords that show the type of the applied grammar rule? | ||||
| 6. Is the exercise contextualized enough to present the usage and the form of the grammar rule? | |||||
| 4. Involvement | 7. Does the exercise grab the interest of students through tackling topics relevant to their background? | ||||
| 8. Does the exercise present the grammar rule appropriately in a graded way (easy to difficult)? | |||||
| 5. Quality of Items | 5.1. Multiple-Choice Items | 9. Are the stems in the multiple-choice items clearly worded and free of ambiguity and irrelevant details? | |||
| 10. Are the alternatives/distractors in the multiple-choice items plausible and clear? | |||||
| 5.2. True/False Items | 11. Are the T/F items worded carefully enough so they can be judged without ambiguity? | ||||
| 12. Do the T/F items include only one main idea? | |||||
| 5.3. Fill-in Items | 13. Is there sufficient context in the fill-in items to convey the intent of the questions? | ||||
| 14. Are the required, correct responses in the fill-in items concise? | |||||
| 5.4. Short-Answer Items | 15. Are the short-answer items formatted so that only one relatively concise answer is possible? | ||||
| 16. Are the short-answer items written in clear and direct questions? | |||||
| Would you like to add any comments? | |||||
Semi-structured interviews
Eight semi-structured interview questions were developed and emailed to participants (see Supplementary Material 3). The interview explored their experiences using Generative AI (GenAI) to create grammar assessments, focusing on the validity, reliability, and contextual relevance of AI-generated test items. Additionally, the participants compared the AI-generated test with the human-made one and discussed the advantages, limitations, and differences between human-generated and AI-generated assessments. The interview questions also examined challenges associated with using GenAI and strategies for optimizing its effectiveness in test creation.
The qualitative data (i.e. participants’ responses) were analyzed using thematic analysis, following the six steps outlined by Braun and Clarke (2006). The thematic analysis was selected because it provides a flexible and rich approach for identifying, analyzing, and reporting patterns within qualitative data. First, the responses were systematically analyzed. Initial codes were then generated inductively, focusing on recurring ideas, phrases, and patterns in how participants described their experiences with AI-generated grammar assessments. These codes were organized and grouped into broader categories, from which potential themes were developed to capture key perceptions, challenges, and suggestions shared by the teachers. The coding process was conducted manually by carefully reading and analyzing the written responses. The codes were reviewed multiple times to identify recurring patterns and refine the codes. This iterative process ensured internal consistency in the development of themes. The themes were reviewed, then clearly defined and labeled based on their relevance to the research questions. The final themes provided insight into how the teachers perceived the usefulness, ease of use, limitations, and practical implications of using AI tools in generating grammar assessments.
Findings
This section presents the data collected through the study’s instruments, beginning with an analysis of the test evaluation checklist, followed by an examination of the interview responses. It highlights key patterns, trends, and insights derived from both quantitative and qualitative data, providing a comprehensive understanding of teachers’ perceptions of AI-generated test items.
Test evaluation checklist
To provide valuable insights into the strengths and limitations of AI-generated and human-made tests, the participants were asked to evaluate both test types separately and rate how well each item met the specified five criteria (i.e. validity, reliability, contextuality, involvement, and quality of items).
The evaluation of AI-generated test
The descriptive analysis in Figure 1 summarizes the participants’ evaluations of the AI-generated test (see Supplementary Material 4). Approximately 66.67% found that the items were aligned with the learning objectives. About 50% reported consistency with the students’ prior learning, while others reported some degree of alignment. All agreed that AI-generated items were suitable for the students’ language proficiency and cognitive levels. The majority, 83.33%, noted that the language of the test was clear, concise, and unambiguous, with 66.67% reporting that the items contained clear keywords of the grammar rule. Half of the participants found the questions were sufficiently contextualized to illustrate the grammar rules’ usage and form, while 41.67% found some contextual relevance within the test items. Approximately 58.33% of respondents reported that the topics were relevant to their students’ backgrounds, and all agreed that the grammar rules were presented appropriately and gradually.
The vertical axis of the vertical bar graph is labeled “Number of Responses” and ranges from 0 to 14 in increments of 2 units. The horizontal axis is labeled “Checklist Questions” and displays sixteen categories from “Question 1” through “Question 16.” The graph contains three colored vertical bars for each checklist question, representing the response options as shown in the legend at the top: Blue bars for “Yes,” Orange bars for “Somewhat,” and Gray bars for “No.” The data from the bars on the graph is as follows: Question 1: Yes: 8, Somewhat: 4, No: Not given. Question 2: Yes: 6, Somewhat: 6, No: Not given. Question 3: Yes: 12, Somewhat: Not given, No: Not given. Question 4: Yes: 1Not given, Somewhat: 2, No: Not given. Question 5: Yes: 8, Somewhat: 4, No: Not given. Question 6: Yes: 6, Somewhat: 5, No: 1. Question 7: Yes: 4, Somewhat: 7, No: Not given. Question 8: Yes: 7, Somewhat: 3, No: 2. Question 9: Yes: 11, Somewhat: Not given, No: 1. Question 10: Yes: 7, Somewhat: 5, No: Not given. Question 11: Yes: 9, Somewhat: 3, No: Not given. Question 12: Yes: 10, Somewhat: 1, No: 1. Question 13: Yes: 1, Somewhat: 8, No: 2. Question 14: Yes: 6, Somewhat: 6, No: Not given. Question 15: Yes: 7, Somewhat: 5, No: Not given. Question 16: Yes: 10, Somewhat: 2, No: Not given.The participants’ evaluation of AI-generated grammar test. Source: Figure created by authors
The vertical axis of the vertical bar graph is labeled “Number of Responses” and ranges from 0 to 14 in increments of 2 units. The horizontal axis is labeled “Checklist Questions” and displays sixteen categories from “Question 1” through “Question 16.” The graph contains three colored vertical bars for each checklist question, representing the response options as shown in the legend at the top: Blue bars for “Yes,” Orange bars for “Somewhat,” and Gray bars for “No.” The data from the bars on the graph is as follows: Question 1: Yes: 8, Somewhat: 4, No: Not given. Question 2: Yes: 6, Somewhat: 6, No: Not given. Question 3: Yes: 12, Somewhat: Not given, No: Not given. Question 4: Yes: 1Not given, Somewhat: 2, No: Not given. Question 5: Yes: 8, Somewhat: 4, No: Not given. Question 6: Yes: 6, Somewhat: 5, No: 1. Question 7: Yes: 4, Somewhat: 7, No: Not given. Question 8: Yes: 7, Somewhat: 3, No: 2. Question 9: Yes: 11, Somewhat: Not given, No: 1. Question 10: Yes: 7, Somewhat: 5, No: Not given. Question 11: Yes: 9, Somewhat: 3, No: Not given. Question 12: Yes: 10, Somewhat: 1, No: 1. Question 13: Yes: 1, Somewhat: 8, No: 2. Question 14: Yes: 6, Somewhat: 6, No: Not given. Question 15: Yes: 7, Somewhat: 5, No: Not given. Question 16: Yes: 10, Somewhat: 2, No: Not given.The participants’ evaluation of AI-generated grammar test. Source: Figure created by authors
Regarding the quality of the items, 91.67% reported that the multiple-choice stems were worded clearly and free of ambiguity. About 58.33% observed that the distractors were plausible and clear, while 41.67% found them “somewhat” clear. Seventy-five percent observed careful wording of True/False items, and 83.33% found that T/F items focused on only one main idea. Regarding fill-in items, 66.67% found the context was “somewhat” adequate, whereas half of the respondents found the correct answers were brief. Approximately 58.33% agreed that short-answer items allowed for a concise response, while 83.33% found them clear and direct in their wording. Overall, the test items generated by AI were viewed positively by the respondents, more so in terms of reliability and clarity of T/F items.
The evaluation of the human-made test
The descriptive analysis in Figure 2 shows the participants’ evaluations of the test created by the human examiner (see Supplementary Material 5). All the participants agreed that the items were aligned with the learning objectives and consistent with the course content. About 75% found the items were appropriate for students’ language proficiency and cognitive levels. Approximately 66.67% noticed that the language of the questions was clear and concise, while almost 83.33% noticed that the questions clearly pointed out the key grammar rules. They all agreed that the questions were contextualized, with about 75% reporting that the items were relevant to students’ background knowledge.
The vertical axis of the vertical bar graph is labeled “Number of Responses” and ranges from 0 to 14 in increments of 2 units. The horizontal axis is labeled “Checklist Questions” and displays sixteen categories from “Question 1” through “Question 16.” The graph contains three colored vertical bars for each checklist question, representing the response options as shown in the legend on the right: Blue bars for “Yes,” Orange bars for “Somewhat,” and Gray bars for “No.” The data from the bars on the graph is as follows: Question 1: Yes: 12, Somewhat: Not given, No: Not given. Question 2: Yes: 12, Somewhat: Not given, No: Not given. Question 3: Yes: 9, Somewhat: 3, No: Not given. Question 4: Yes: 8, Somewhat: 4, No: Not given. Question 5: Yes: 10, Somewhat: 2, No: Not given. Question 6: Yes: 12, Somewhat: Not given, No: Not given. Question 7: Yes: 9, Somewhat: 3, No: Not given. Question 8: Yes: 9, Somewhat: 3, No: Not given. Question 9: Yes: 11, Somewhat: Not given, No: 1. Question 10: Yes: 10, Somewhat: 1, No: 1. Question 11: Yes: 9, Somewhat: 3, No: Not given. Question 12: Yes: 8, Somewhat: 3, No: 1. Question 13: Yes: 10, Somewhat: 1, No: Not given. Question 14: Yes: 7, Somewhat: 4, No: Not given. Question 15: Yes: 11, Somewhat: 1, No: Not given. Question 16: Yes: 12, Somewhat: Not given, No: Not given.The participants’ evaluation of the human-made grammar test. Source: Figure created by authors
The vertical axis of the vertical bar graph is labeled “Number of Responses” and ranges from 0 to 14 in increments of 2 units. The horizontal axis is labeled “Checklist Questions” and displays sixteen categories from “Question 1” through “Question 16.” The graph contains three colored vertical bars for each checklist question, representing the response options as shown in the legend on the right: Blue bars for “Yes,” Orange bars for “Somewhat,” and Gray bars for “No.” The data from the bars on the graph is as follows: Question 1: Yes: 12, Somewhat: Not given, No: Not given. Question 2: Yes: 12, Somewhat: Not given, No: Not given. Question 3: Yes: 9, Somewhat: 3, No: Not given. Question 4: Yes: 8, Somewhat: 4, No: Not given. Question 5: Yes: 10, Somewhat: 2, No: Not given. Question 6: Yes: 12, Somewhat: Not given, No: Not given. Question 7: Yes: 9, Somewhat: 3, No: Not given. Question 8: Yes: 9, Somewhat: 3, No: Not given. Question 9: Yes: 11, Somewhat: Not given, No: 1. Question 10: Yes: 10, Somewhat: 1, No: 1. Question 11: Yes: 9, Somewhat: 3, No: Not given. Question 12: Yes: 8, Somewhat: 3, No: 1. Question 13: Yes: 10, Somewhat: 1, No: Not given. Question 14: Yes: 7, Somewhat: 4, No: Not given. Question 15: Yes: 11, Somewhat: 1, No: Not given. Question 16: Yes: 12, Somewhat: Not given, No: Not given.The participants’ evaluation of the human-made grammar test. Source: Figure created by authors
The majority, 91.67%, agreed that the multiple-choice stems were clear and unambiguous, and 83.33% found the distractors plausible and clear. As far as True/False items, about 75% noticed the items were worded carefully, and 66.67% observed that each T/F item covered one main idea at a time. In the case of fill-in questions, 83.33% found the context provided was adequate, while 58.33% reported that the correct responses were concise enough. Short-answer items were praised by the majority, 91.67%, for clarity and concise responses. Hence, the test created by the human examiner was viewed positively by the respondents in terms of validity, contextuality, involvement, and across various types of items (i.e. MCQs, fill-in-the-blanks, and short answer items).
AI-generated test vs human-made test
As shown in Figure 3, the evaluation of AI-generated and human-made tests revealed significant differences. On the aspect of validity, the human-made test outperformed the AI-generated test by scoring 100% on alignment to learning objectives and consistency with course content, against 66.67% and 50%, respectively, for the AI-generated test. As far as reliability, AI-generated items scored higher in terms of items’ appropriateness to students’ language proficiency and cognitive levels, with a perfect score of 100%, while the human-made test scored 75%. Additionally, the AI-generated test scored 83.33% for language clarity and conciseness, as opposed to the human-made test with 66.67%. The human-made test demonstrated superior contextuality, achieving 83.33% in integrating the grammar rule and 100% in presentation. In contrast, the AI-generated test scored lower, with 66.67% for grammar rule integration and 50% for presentation.
The vertical axis of the vertical bar graph ranges from 0.00 percent to 120.00 percent in increments of 20 percent. The horizontal axis is labeled with sixteen categories: “Validity 1,” “Validity 2,” “Reliability 1,” “Reliability 2,” “Contextuality 1,” “Contextuality 2,” “Involvement 1,” “Involvement 2,” “Quality of M C Q s 1,” “Quality of M C Q s 2,” “Quality of T over F 1,” “Quality of T over F 2,” “Quality of Fill-in 1,” “Quality of Fill-in 2,” “Quality of Short-Answer 1,” and “Quality of Short-Answer 2.” The graph contA Ins two colored vertical bars for each category, representing the test type as shown in the legend: Blue bars for “The quality of A I generated test items” and Orange bars for “The quality of human made test.” The data from the bars on the graph is as follows: Validity 1: A I: 66.7 percent, Human: 100 percent. Validity 2: A I: 50.17 percent, Human: 100 percent. Reliability 1: A I: 100 percent, Human: 75.08 percent. Reliability 2: A I: 82.69 percent, Human: 66.43 percent. Contextuality 1: A I: 66.43 percent, Human: 82.69 percent. Contextuality 2: A I: 49.32 percent, Human: 100 percent. Involvement 1: A I: 32.87 percent, Human: 74.74 percent. Involvement 2: A I: 57.78 percent, Human: 74.74 percent. Quality of M C Q s 1: A I: 91.34 percent, Human: 91.34 percent. Quality of M C Q s 2: A I: 58.13 percent, Human: 83.39 percent. Quality of T over F 1: A I: 75.08 percent, Human: 75.08 percent. Quality of T over F 2: A I: 82.35 percent, Human: 66.09 percent. Quality of Fill-in 1: A I: 16.95 percent, Human: 83.73 percent. Quality of Fill-in 2: A I: 49.82 percent, Human: 57.78 percent. Quality of Short-Answer 1: A I: 58.47 percent, Human: 91.00 percent. Quality of Short-Answer 2: A I: 83.39 percent, Human: 100 percent. Note: All numerical values are approximated.AI-generated vs human-made tests. Source: Figure created by authors
The vertical axis of the vertical bar graph ranges from 0.00 percent to 120.00 percent in increments of 20 percent. The horizontal axis is labeled with sixteen categories: “Validity 1,” “Validity 2,” “Reliability 1,” “Reliability 2,” “Contextuality 1,” “Contextuality 2,” “Involvement 1,” “Involvement 2,” “Quality of M C Q s 1,” “Quality of M C Q s 2,” “Quality of T over F 1,” “Quality of T over F 2,” “Quality of Fill-in 1,” “Quality of Fill-in 2,” “Quality of Short-Answer 1,” and “Quality of Short-Answer 2.” The graph contA Ins two colored vertical bars for each category, representing the test type as shown in the legend: Blue bars for “The quality of A I generated test items” and Orange bars for “The quality of human made test.” The data from the bars on the graph is as follows: Validity 1: A I: 66.7 percent, Human: 100 percent. Validity 2: A I: 50.17 percent, Human: 100 percent. Reliability 1: A I: 100 percent, Human: 75.08 percent. Reliability 2: A I: 82.69 percent, Human: 66.43 percent. Contextuality 1: A I: 66.43 percent, Human: 82.69 percent. Contextuality 2: A I: 49.32 percent, Human: 100 percent. Involvement 1: A I: 32.87 percent, Human: 74.74 percent. Involvement 2: A I: 57.78 percent, Human: 74.74 percent. Quality of M C Q s 1: A I: 91.34 percent, Human: 91.34 percent. Quality of M C Q s 2: A I: 58.13 percent, Human: 83.39 percent. Quality of T over F 1: A I: 75.08 percent, Human: 75.08 percent. Quality of T over F 2: A I: 82.35 percent, Human: 66.09 percent. Quality of Fill-in 1: A I: 16.95 percent, Human: 83.73 percent. Quality of Fill-in 2: A I: 49.82 percent, Human: 57.78 percent. Quality of Short-Answer 1: A I: 58.47 percent, Human: 91.00 percent. Quality of Short-Answer 2: A I: 83.39 percent, Human: 100 percent. Note: All numerical values are approximated.AI-generated vs human-made tests. Source: Figure created by authors
In terms of involvement, the human-made test scored higher, achieving 75% for relevance to students’ backgrounds and appropriate presentation of grammar rules. On the other hand, the AI-generated test received lower scores, with 33.33% for relevance and 58.33% for grammar rule presentation. The two test types equally performed high in terms of clarity of multiple-choice stems at 91.67%, but the distractors in the human-made test were more plausible at 83.33%, compared to 58.33% for the AI-generated test. As far as True/False items, clarity was rated at 75% for both test types, but the AI-generated test was slightly higher in terms of T/F items’ addressing one main idea at a time, 83.33% vs 66.67%. Regarding the fill-in items, significant differences were observed, with the human-made test scoring 83.33% on contextual clarity, while the AI-generated test scored just 16.66%. In terms of conciseness, it scored a bit higher at 58% for the human-made test and 50% for the AI-generated test. For the case of short-answer items, the human-made test scored 91% for formatting and 100% for clarity, while the AI-generated test achieved 58.33% on both criteria.
To summarize, the AI-generated test received an average rating of 63.33%, whereas the human-made test scored significantly higher, with a mean of 83.33%. These scores reinforce the above findings, which reveal that AI-generated assessments show promise, especially in terms of reliability and clarity. However, human-made tests are still perceived as more effective and pedagogically aligned in the context of grammar evaluation.
Interview responses analysis
After receiving the participants’ responses to the interview questions, the researchers started the analysis process to address the research questions. The analysis style conducted is thematic, in which the researchers reviewed the responses repeatedly to extract codes. These codes were then used to develop emerging themes. The emerging themes included benefits, challenges, and actions to be considered. There were three codes in terms of the benefits theme: saving time, resources/questions diversity or variety, and personalization/learning styles. The challenges’ themes included context clarity/ambiguity, prompts/parameters, over-reliance, ethical issues, and errors. There were only two codes relevant to the actions-to-be-considered theme: monitor, human editing/interference/modifications/review (see Supplementary Material 6).
All the participants assured that GenAI serves as an important timesaver, generating many questions instantaneously, minimizing teachers’ efforts, thereby positively impacting their satisfaction and burnout level (see Extracts Group 1, P4 & P8). This is attributed to the substantial workload imposed on university teachers, which involves administrative tasks, lecture preparations, constant follow-up with students, preparing assessments, and conducting research. They also highlighted the important role of human intervention through editing or modifying the generated questions to align with the course objectives, cultural context, learning styles, and learners’ interests (see Extracts Group 1, P5 & P10), as human interference in the assessment process is crucial and inevitable to reach a higher level of assessment practicality, reliability, and validity.
Extracts Group 1
“AI tools like ChatGPT can quickly generate a variety of questions and formats, which saves a lot of time.” (P4)
“It saves time as AI can generate large number of questions quickly and efficiently.” (P8)
“The contexts provided can be bland, irrelevant, or typical.” (P5)
“The exercises still need to be revised and amended by the instructor to make sure that they are aligned with course objectives and content.” (P10)
Another benefit of utilizing GenAI in generating grammar quizzes and exams is the diversity of questions (see Extracts Group 2, P3 & P4). AI can generate various types and forms of grammar questions, including gap fill, multiple-choice, true-or-false questions, and more. Nevertheless, this requires accurate prompts (see Extracts Group 2, P1 & P11). The specificity of a teacher’s request is essential and conducive to accurate results, not only for grammar assessments but also for almost every request from GenAI. A teacher, for instance, must specify the course objectives, types of questions, cultural context, learners’ age, level of proficiency, and similar prompts.
Extracts Group 2
“AI has unlimited resources that enables it to give diverse choices.” (P3)
“AI tools can generate diverse question types and formats, offering a range of difficulty levels suitable for different learners.” (P4)
“It needs guided prompts to be more effective if language teacher is not aware of how to ask the questions in a proper and effective way, the response will not be supportive.” (P1)
“GenAI can generate grammar exams that are well-aligned with course objectives, but only if provided with the right prompts and information (the specific learning goals) upfront” (P11)
A third benefit is the personalization of assessment (see Extracts Group 3, P2 & P4, 1). Personalization is a teaching model that tailors the learning experience to the student’s learning styles, existing knowledge, skills, and personal interests, contrasting with the traditional one-size-fits-all approach (Morin, 2019). Still, there is a concern that GenAI may sometimes struggle to address this point due to inaccurate prompting, context ambiguity, or errors; such an occurrence requires attention through monitoring (see Extracts Group 3, P9 & P10). Nevertheless, this calls teachers to be entirely attentive to such errors, as GenAI is not impeccable. Additionally, GenAI errors may manifest in content, grammar rules, or sentence structures eventually conducing to distorted instruction or assessment (see Extracts Group 3, P4, 2 & P9, 2). That is why human interference is indispensable.
Extracts Group 3
“It will personalize exercises” (P2)
“GenAI can also tailor exercises to the specific needs and proficiency levels of individual learners, enhancing personalized learning.” (P4, 1)
“Monitor the effectiveness of AI-generated exercises and exams through students’ performance and feedback to alter and adjust the items if needed.” (P9)
“The course coordinator and instructors must monitor any limitations or insufficiency faced while using those questions.” (P10)
“I think errors may happen as well as ambiguities when using AI in generating materials.” (P3)
“Ensuring that the generated content is free from errors and aligns accurately with curriculum standards can be difficult.” (P4, 2)
Discussion
What are the differences between AI-generated and human-made tests?
The analysis comparing AI-generated and human-made tests reveals several key findings.
Validity: The human-made test scored higher in terms of aligning with learning objectives and maintaining consistency with the course content. Owan et al. (2023) highlighted the crucial role of teachers in designing and contextualizing assessments; they can create assessments that are specifically tailored to meet students’ needs due to their deep understanding of the content and learning objectives.
Reliability: On the other hand, AI-generated tests scored higher in terms of reliability, particularly in language clarity and appropriateness for students’ cognitive levels. This finding is supported by research, indicating that AI can create items with precise and straightforward language to ensure clarity and reduce ambiguity that might be present in human-written questions. Additionally, it has been found that AI can tailor assessments based on students’ proficiency levels, ensuring that each student is challenged appropriately (Stоšić and Malyuga, 2024).
Contextuality and Involvement: The human-made test excelled in these criteria. Owan et al. (2023) stated that teachers can enrich tests with relevant items due to their nuanced understanding and expertise, thereby boosting students’ engagement and enhancing the overall learning experience. Although AI has been found efficient and supportive, the expertise and contextual knowledge of teachers are indispensable for creating engaging and meaningful assessments. As noted by Stоšić and Malyuga (2024), “[AI] may struggle with understanding context, leading to inaccurate assessments of language skills in certain situations” (p. 29). Additionally, Chan and Louisa (2023) recognized the unique qualities of human experts, that is, their ability “… to provide real-world examples and experiences that AI systems may not be able to offer” (p. 4). On the contrary, AI may struggle with contextualization due to its limitations in understanding language nuances and cultural contexts that human examiners can naturally comprehend (Abida et al., 2023; Stоšić and Malyuga, 2024; Kic-Drgas and Kılıçkaya, 2024). According to Abida et al. (2023), “[Chatbots] may not yet fully capture differences in the use of more subtle language elements, such as idioms, slang, or cultural references” (p. 143).
The Quality of Items: Both AI-generated and human-made tests scored the same in terms of the clarity of multiple-choice stems. However, they differed in the plausibility of distractors, as AI might occasionally fail to generate distractors that address the complex elements of the language (Abida et al., 2023; Kic-Drgas and Kılıçkaya, 2024; Owan et al., 2023; Stоšić and Malyuga, 2024). Overall, the human-made test provided clear and contextually rich questions across the various test items (e.g. multiple-choice questions, true/false statements, fill-in-the-blanks, and short-answer questions). Malec (2024) concluded, “Although the use of artificial intelligence has an unquestionably positive impact on test practicality, ChatGPT-generated multiple-choice items cannot yet be used in operational settings without human moderation” (p. 836). Nonetheless, the AI-generated test scored higher in crafting true/false questions due to its algorithms that can eliminate ambiguities and, in turn, enhance clarity and precision. According to O (2024), AI was very useful for crafting true/false items in her study. For instance, when AI was provided with a sample true statement and asked to generate multiple true statements, it successfully produced statements that closely resembled the sample. It implies that AI cannot only replicate but also create new and high-quality true/false items.
Based on these findings, a collaborative approach between AI and human examiners can optimize language assessment practices. According to the present study, the human-made test excelled in validity, contextuality, and involvement criteria due to the human examiner’s expertise in aligning assessments with learning objectives and integrating questions relevant to the context. On the other hand, the AI-generated test demonstrated strengths in reliability and clarity of language, as algorithms can ensure consistency and eliminate ambiguities within test items. Chan and Louisa (2023) noted that despite AI’s capabilities, AI is viewed as “… [cognitive prostheses] that can aid teaching and learning, but not yet capable of replacing the values of human thoughts” (p. 4). Thus, recognizing the role of human experts alongside AI is key in language assessment. AI’s ability to craft certain items, such as true/false, is evident. However, its lack of contextual understanding highlights the critical role of human examiners and their indispensable knowledge and insights. Thao (2023) affirmed that, “AI will support teachers in the way they want; however, it cannot replace them” (p.114).
What are the potential benefits of using AI-generated grammar exercises or tests in language classrooms?
Based on the reported results and interview analysis, AI offers three main benefits when used to generate grammar exercises, quizzes, or tests. The first is its potential to save time (Barenkamp et al., 2020), which can significantly support university language instructors by easing the test creation process and allowing them to allocate more time to other instructional or administrative duties. Secondly, it provides various types of questions, including gap fills, multiple-choice questions, and true/false (Fitria, 2021). Such variety enables teachers to design more comprehensive exams that assess various aspects of grammar. Furthermore, it makes the assessment process more engaging and less predictable, thereby minimizing rote memorization and encouraging deeper learning. The final benefit is personalization, which fosters students’ engagement, boosts self-efficacy, improves academic outcomes, and enhances retention (Pane et al., 2017). Therefore, with the help of AI, grammar learning can be more personalized by addressing the needs of learners and generating assessments aligned with the latest teaching approaches and learning objectives.
What are the challenges reported by university English language teachers regarding using AI-generated grammar activities or tests?
The participants reported five challenges of using AI in creating grammar assessments, which are consistent with concerns raised in prior research. First, AI, as a machine, cannot fully comprehend the complexities of human nature or societal changes, as it only processes algorithms and mathematical calculations. Only human examiners can effectively address cultural, educational, and cognitive contexts (Abida et al., 2023; Stоšić and Malyuga, 2024; Chan and Louisa, 2023). Second, the effectiveness of AI-generated content heavily relies on the specificity of the prompts used, as vague instructions may lead to flawed or irrelevant output (Crompton et al., 2024; Lalira et al., 2024). Since AI can analyze vast amounts of data, being specific is crucial. Thus, when making a request, the teacher must provide information about the targeted group of learners, course objectives, expected outcomes, and perhaps templates for the desired grammar questions. Only then can AI provide the expected results.
Also, reliance on technology in education has been exacerbated since the emergence of COVID-19. Online learning approaches have gained widespread acceptance since then; however, in the age of AI, there has been an over-reliance on technology. In this sense, utilizing technological tools in language education does not mean entirely relying on AI in creating learning resources, as it will negatively impact teachers’ innovation, critical thinking, and problem-solving skills (Thao, 2023), In addition, errors and ethical issues are also pitfalls of using AI in language education. Although AI is thought to be accurate, it may commit language errors when generating grammar assessments, which leads to distorted instructions if not carefully reviewed by human examiners (Kic-Drgas and Kılıçkaya, 2024).
What are some practical strategies to help language instructors tackle the limitations of AI-generated exercises?
Regular human review is essential to validate the accuracy, appropriateness, and pedagogical fit of AI-generated tasks. Rather than replacing instructors, AI should be viewed as a cognitive assistant that supports teaching without eliminating the role of human insight (Chan and Louisa, 2023). Incorporating students’ feedback into the evaluation of AI-generated assessments also enhances their relevance and alignment with learner needs, reinforcing principles of personalized learning (Pane et al., 2017).
Furthermore, the findings carry important implications for teacher training and institutional policies. Teacher education programs should incorporate modules on AI literacy, equipping future teachers with the skills to critically evaluate, adapt, and integrate AI-generated content into their curriculum. This includes training teachers not only to use AI tools but also to understand their limitations, ethical considerations, and the pedagogical reasoning needed to refine AI-generated assessments to suit diverse learners. Additionally, institutional policies should promote professional development workshops that emphasize collaborative AI-human teaching approaches, ensuring that AI is leveraged meaningfully rather than uncritically adopted.
Conclusion
This study investigated university English language teachers’ perceptions of using AI tools to generate grammar assessments, highlighting both the advantages and limitations of this approach. While AI-generated assessments offer time-saving benefits, diverse question formats, and personalization opportunities, they often fall short in contextual relevance and learners’ engagement when compared to human-made assessments. Human-generated tests were found to better align with instructional goals, whereas AI tools demonstrated higher consistency and linguistic clarity. To ensure effective integration of AI in language assessment, the study recommends two key strategies: (1) human review and adaptation of AI-generated materials to fit learners’ needs, and (2) continuous evaluation of AI tools.
While the study provides valuable insights into the use of AI in grammar assessment, it is not without limitations. These include a relatively small number of participating teachers and a lack of extensive prior research specifically addressing AI applications in grammar teaching and assessment. These limitations highlight the need for further research in this area. Additionally, institutions should consider developing clear policies that regulate the ethical use of AI in assessment and support teacher training programs that equip instructors with the skills needed to critically evaluate and effectively use AI-generated content. By combining technological efficiency with human pedagogical insight, language programs can maximize the benefits of AI while maintaining the quality and integrity of assessments.
The supplementary material for this article can be found online.

