Наука и инновации 3-jild vation-son (2025) · 35–38-betlar
ACHIEVING HIGHER ACCURACY IN CLASSIFYING UZBEK WORDS INTO GRAMMATICAL CATEGORIES USING THE CRF MODEL
Kobilov, Sami, Nazarov, Javohir, Rabbimov, Ilyos
DOI: 10.5281/zenodo.17189243 · Manbada o'qish →
Annotatsiya
The classification of words into grammatical categories (part-of-speech tagging) is a fundamental task in natural language processing (NLP). For morphologically rich languages such as Uzbek, this process becomes more challenging due to complex affixation, agglutinative word forms, and limited resources. This paper investigates the application of the Conditional Random Fields (CRF) model for Uzbek word classification, aiming to achieve higher accuracy compared to traditional approaches such as Hidden Markov Models (HMM). By incorporating contextual and morphological features, CRF achieves more reliable tagging. Several Uzbek sentence examples are analyzed, with CRF applied step by step to demonstrate its advantages. Experimental results show that CRF significantly improves accuracy, achieving 92.7% compared to 84.3% for HMM.
Uzbek language, part-of-speech tagging, word classification, CRF model, natural language processing.
Metadata manbasi: jurnal OAI-PMH arxivi · Sindex to'liq matnni saqlamaydi, manbaga havola beradi.