Publication Details
Issue: Vol 7, No 1 (2026)
Pages: 269-272
ISSN: 2660-6828

Abstract

This article analyzes the challenges of automatic part-of-speech identification (PoS tagging) in the Uzbek language, existing approaches, and the possibilities of using practical tools. Due to the agglutinative nature of Uzbek, PoS tagging requires careful consideration of morphological analysis, contextual meaning, and variations in affixal forms. The paper discusses the effectiveness of rule-based and statistical PoS taggers, particularly those developed using the Hidden Markov Model (HMM), as well as the advantages of the BBPOS system based on neural networks. In addition, the article demonstrates how morphological analysis results obtained through the uznatcorpora.uz platform provide a solid foundation for PoS tagging in Uzbek. The research findings highlight the necessity of creating PoS-tagged texts for an Uzbek–English parallel corpus and reveal the linguistic and practical value of such a corpus.

Keywords
PoS tagging Uzbek language corpus linguistics morphological analysis Hidden Markov Model BERT uznatcorpora.uz parallel corpus linguistic annotation