Рақамли технологияларнинг назарий ва амалий масалалари 5-jild 3-son (2023) · 47-56-betlar
An Integrated Analysis of Multilingual Texts Spanning Dual Alphabets
Адилова, Ф.Т., Давронов, Р.Р., Сафаров, Р.А.
Annotatsiya
Language recognition in natural language processing (NLP) aims to determine the specific language of a text or document. As the number of languages increases, this task becomes more complex. This study introduces a detailed model for detecting languages from text, with an emphasis on the Latin-Cyrillic script of the Uzbek language. Noting the research gap in this domain, we unveil a precise Uzbek Latin-Cyrillic script recognition model leveraging an apt transformer architecture. The model was tested on our self-compiled Uzbek language corpus, which also offers a robust benchmark for subsequent Uzbek language identification studies. Our approach encompasses 21 languages, including Uzbek, across both Latin and Cyrillic alphabets. Our findings highlight that the XLM-RoBERTa transformer-driven language detection model significantly outperforms its predecessors in terms of accuracy and efficiency.
NLPMultilingual Language ModelsCloud Natural Language APIOpen AIChatGPTmodel compressiontransformerNLPМногоязычные языковые моделиОблачный API естественного языка
Metadata manbasi: jurnal OAI-PMH arxivi · Sindex toʻliq matnni saqlamaydi, manbaga havola beradi.