Рақамли технологияларнинг назарий ва амалий масалалари 9-том 3-сан (2026) · 155-165-беттер
Comparative analysis of TF-IDF and word embedding methods for text classification in the Uzbek language
Рахимов, Р.Т.
Аннотация
In this work, a comparative analysis of vectorization methods and classification algorithms in the classification of text data in the Uzbek language was carried out. In the study, a new dataset called “UzText-600” was developed by the author, consisting of 600 texts collected from various sources in the Uzbek language. The dataset is divided into 6 thematic classes: automobile, international events, culture and art, sports, history and heritage, politics and economics. Five vectorization methods (TF-IDF, Word2Vec, FastText, Doc2Vec, BERT) and 10 classifiers (traditional: Support vector machine (SVM), Logistic regression, Naive Bayes; ensemble: Random forest, XGBoost, LightGBM; deep learning: Long short-term memory (LSTM), BidirectionalLSTM (BiLSTM), Transformer) were applied to this dataset. Experimental results showed that the combination of BERT + XGBoost achieved the highest result with an accuracy of 0.942, while TF-IDF + SVM recorded the best performance among traditional methods with an accuracy of 0.865. Statistical analyses (paired t-test, ANOVA) proved that the methods based on Bidirectional encoder representations from transformers (BERT) were significantly superior to other methods. This work provides an initial benchmark dataset specifically created for the Uzbek language and serves as a basis for future research. All results were analyzed in detail through graphs and tables.
узбекский языкклассификация текстаTF-IDFWord2VecFastTextBERTXGBoostTransformerсравнительный анализUzText-600
Метадайындар булагы: журналдын OAI-PMH архиви · Sindex толук текстти сактабайт, булакка шилтеме берет.