Рақамли технологияларнинг назарий ва амалий масалалари Ҷилди 5 № 3 (2023) · Саҳифаҳои 47-56
An Integrated Analysis of Multilingual Texts Spanning Dual Alphabets
Адилова, Ф.Т., Давронов, Р.Р., Сафаров, Р.А.
Аннотатсия
Language recognition in natural language processing (NLP) aims to determine the specific language of a text or document. As the number of languages increases, this task becomes more complex. This study introduces a detailed model for detecting languages from text, with an emphasis on the Latin-Cyrillic script of the Uzbek language. Noting the research gap in this domain, we unveil a precise Uzbek Latin-Cyrillic script recognition model leveraging an apt transformer architecture. The model was tested on our self-compiled Uzbek language corpus, which also offers a robust benchmark for subsequent Uzbek language identification studies. Our approach encompasses 21 languages, including Uzbek, across both Latin and Cyrillic alphabets. Our findings highlight that the XLM-RoBERTa transformer-driven language detection model significantly outperforms its predecessors in terms of accuracy and efficiency.
NLPMultilingual Language ModelsCloud Natural Language APIOpen AIChatGPTmodel compressiontransformerNLPМногоязычные языковые моделиОблачный API естественного языка
Манбаи метамаълумот: бойгонии OAI-PMH-и маҷалла · Sindex матни пурраро нигоҳ намедорад, ба манбаъ пайванд медиҳад.