Рақамли технологияларнинг назарий ва амалий масалалари 5-tom 3-san (2023) · 47-56-betler
An Integrated Analysis of Multilingual Texts Spanning Dual Alphabets
Адилова, Ф.Т., Давронов, Р.Р., Сафаров, Р.А.
Annotaciya
Language recognition in natural language processing (NLP) aims to determine the specific language of a text or document. As the number of languages increases, this task becomes more complex. This study introduces a detailed model for detecting languages from text, with an emphasis on the Latin-Cyrillic script of the Uzbek language. Noting the research gap in this domain, we unveil a precise Uzbek Latin-Cyrillic script recognition model leveraging an apt transformer architecture. The model was tested on our self-compiled Uzbek language corpus, which also offers a robust benchmark for subsequent Uzbek language identification studies. Our approach encompasses 21 languages, including Uzbek, across both Latin and Cyrillic alphabets. Our findings highlight that the XLM-RoBERTa transformer-driven language detection model significantly outperforms its predecessors in terms of accuracy and efficiency.
NLPMultilingual Language ModelsCloud Natural Language APIOpen AIChatGPTmodel compressiontransformerМногоязычные языковые моделиОблачный API естественного языкаОткрытый ИИ
Metadata derekkózi: jurnal OAI-PMH arxivi · Sindex tolıq mátindi saqlamaydı, derekkózge silteme beredi.