CORPUS-BASED EVALUATION OF AI WRITING FEEDBACK TOOLS' ACCURACY AGAINST HUMAN-ANNOTATED LEARNER CORPORA
Abstract
The use of automated writing evaluation (AWE) tools and generative artificial intelligence (AI) writing assistants like Grammarly, Microsoft Editor, and large language model (LLM) based systems of L2 writing feedback has been widely adopted in second and foreign language (L2) writing pedagogy. Though these tools are popular, doubts linger over their ability to detect and correct the error types human raters can identify in authentic learner texts using the help of existing annotation schemes. The aim of this study is to perform a corpus-based assessment of the accuracy of the three most popular AI writing feedback tools against two human annotated learner corpora, namely the Cambridge Learner Corpus First Certificate in English (FCE) and the NUS Corpus of Learner English (NUCLE). The study uses quantitative, comparative research design based on Error Analysis and Automated Writing Evaluation validity framework, which calculate precision, recall, and F0.5 scores for each tool for every of seven error categories, namely error category of article, determiner, preposition, verbs' tense and agreement, spelling, punctuation, word choice and word order. Three hundred error-annotated sentences were sampled and AI-made corrections were compared systematically against the gold standard human corrections by the MaxMatch scoring method. Results show that AI tools can accurately detect mechanical errors (spelling and punctuation) and provide high precision and recall on these types of error, but their performance is significantly lower on the context-dependent types of errors (word choice and usage of articles, word order). This suggests that there is an important difference between the ability to detect errors at the surface level and the ability to make judgements about errors at the level of discourse, which is what expert human feedback is able to do. The study ends with pedagogical and developmental suggestions for effective incorporation of AI feedback tools in L2 writing classrooms
