Department of Computer Science, Government College (Autonomous), Rajahmundry, Andhra Pradesh, India.
Global Journal of Engineering and Technology Advances, 2026, 27(01), 139-151
Article DOI: 10.30574/gjeta.2026.27.1.0093
Received on 13 March 2026; revised on 21 April 2026; accepted on 23 April 2026
The rapid expansion of digital communication and Large Language Models (LLMs) has led to an unprecedented surge in machine-generated content, raising significant concerns regarding academic integrity, misinformation, and content authenticity. Distinguishing between human-written and AI-generated text manually is increasingly difficult due to the sophisticated nature of modern AI outputs. This study proposes a deep learning-based framework to automatically classify textual data into two categories: AI-generated and human-written.
The research utilised a comprehensive dataset of 28,747 records, which was subjected to a rigorous preprocessing pipeline including data cleaning, tokenisation, stop-word removal, and stemming to ensure high data quality. Exploratory Data Analysis (EDA) was performed to uncover structural differences, revealing that while human text often exhibits a "long tail" in word distribution, AI-generated content follows more constrained and uniform patterns.
A Bidirectional Long Short-Term Memory (BILSTM) model was implemented for the classification task. By processing text in both forward and backward directions, the BILSTM architecture captures deep contextual dependencies and semantic relationships that traditional machine learning models often overlook. The model's performance was evaluated using standard metrics, including accuracy, precision, recall, and F1-score.
The proposed BILSTM model achieved a high-test accuracy of 94.96%, outperforming the standard LSTM model and demonstrating exceptional reliability in identifying text origins. These results highlight the effectiveness of bidirectional deep learning architectures in handling complex natural language sequences. This system holds significant potential for real-world applications in academic plagiarism detection, online content moderation, and AI content verification. Future work includes extending the model to support multilingual datasets and integrating Transformer-based architectures like BERT for enhanced linguistic understanding.
AI-Generated Text Detection; Natural Language Processing (NLP); Deep Learning; Bidirectional LSTM (BILSTM); LSTM; Text Classification; Human vs. AI Text; Word Embedding; Sequence Padding; Content Authenticity
Preview Article PDF
Ullamparthi Durga Venkata Santhosh and Suneel Kumar Duvvuri. Content authenticity detection using bidirectional LSTM networks. Global Journal of Engineering and Technology Advances, 2026, 27(01), 139-151. Article DOI: https://doi.org/10.30574/gjeta.2026.27.1.0093.





