Innovation science and technologiy Volume 2 Issue 6 (2026)
FUNDAMENTAL METHODS FOR IDENTIFYING AI-GENERATED DATA IN ACADEMIC AND ONLINE PUBLICATIONS USING MACHINE LEARNING AND NLP MODELS
Abdullaev, Munis, Kungratov, Ilmurod
Abstract
This article provides a comprehensive scientific analysis of the fundamental and practical methodologies for identifying textsgenerated by Artificial Intelligence (AI) within academic and digital publishing ecosystems. The rapid maturation of generative languagearchitectures is fundamentally transforming traditional copyright paradigms and the principles of academic integrity. The primaryobjective of this research is to develop, test, and evaluate innovative methodologies for distinguishing synthetic data using NaturalLanguage Processing (NLP) and Machine Learning (ML) classification models. The study comparatively evaluates the effectivenessof stylometric feature extraction, zero-shot probability distribution analysis, and transformer-based deep learning classifiers. Empiricalresults confirm that traditional plagiarism systems based on exact lexical matching have completely lost their functional viability.Concurrently, the hybrid-ensemble architecture proposed in this study demonstrated high resilience against complex adversarialevasion attacks. The research findings serve as a critical guide for higher education institutions and scientific journals to optimize theirverification mechanisms and ensure academic honesty
Synthetic, text, generative, architecture, natural, language, processing, transformer, academic, integrity, stylometry, verification, classifier, probability, algorithm
Metadata source: the journal's OAI-PMH archive · Sindex does not store the full text; it links to the source.