Native language identification from text using a fine-tuned GPT-2 model

Native language identification (NLI) is a critical task in computational linguistics, supporting applications such as personalized language learning, forensic analysis, and machine translation. This study investigates the use of a fine-tuned GPT-2 model to enhance NLI accuracy. Using the NLI-PT data...

Full description

Saved in:

Bibliographic Details
Main Author:	Yuzhe Nie
Format:	Article
Language:	English
Published:	PeerJ Inc. 2025-05-01
Series:	PeerJ Computer Science
Subjects:	Native language identification ChatGPT NLP Deep learning
Online Access:	https://peerj.com/articles/cs-2909.pdf
Tags:	Add Tag No Tags, Be the first to tag this record!

Description
Summary:	Native language identification (NLI) is a critical task in computational linguistics, supporting applications such as personalized language learning, forensic analysis, and machine translation. This study investigates the use of a fine-tuned GPT-2 model to enhance NLI accuracy. Using the NLI-PT dataset, we preprocess and fine-tune GPT-2 to classify the native language of learners based on their Portuguese-written texts. Our approach leverages deep learning techniques, including tokenization, embedding extraction, and multi-layer transformer-based classification. Experimental results show that our fine-tuned GPT-2 model significantly outperforms traditional machine learning methods (e.g., SVM, Random Forest) and other pre-trained language models (e.g., BERT, RoBERTa, BioBERT), achieving a weighted F1 score of 0.9419 and an accuracy of 94.65%. These results show that large transformer models work well for native language identification and can help guide future research in personalized language tools and artificial intelligence (AI)-based education.
ISSN:	2376-5992

Native language identification from text using a fine-tuned GPT-2 model

Similar Items