Publication Details
Issue: Vol 7, No 1 (2025)
Pages: 32-38
ISSN: 2660-6828

Abstract

The advancement of corpus linguistics in Uzbekistan requires the establishment of efficient digital tools for linguistic data retrieval and analysis. Despite the creation of several electronic corpora, comprehensive studies on search mechanisms—specifically sorting, storing, and lexical filtering in the Uzbek language corpus—remain limited. Current corpus systems lack a detailed methodological framework for organizing search results, handling lexical parameters such as lemmas and morphemes, and ensuring user-friendly data export and management. This study aims to analyze and systematize search mechanisms for the Uzbek language corpus by focusing on sorting, storing, and lexical search parameters, and adapting international corpus practices to the morphological complexity of Uzbek. The insights gained through the findings also unveil how even the very data that is presented within some of these resources can be made even more meaningful and discoverable through combining alphabetical, frequency and metadata-based sorting with lemma- and morpheme-based search capabilities to improve search functionality as a whole. Moreover, exporting (CSV, XML, JSON) and history-saving functions make sure that the software will be usable in the long-term for research. The research presents a general model for combining computational and linguistic principles to increase the efficiency of corpus search and an adaptive model of dealing with agglutinative structures. The proposed system strengthens the methodological foundation of Uzbek corpus linguistics, facilitates corpus-based research and teaching, and supports the development of computational linguistics in Uzbekistan by transforming the Uzbek corpus into an interactive, analytical, and educational digital resource.

Keywords
corpus of the Uzbek language search engine lemma morpheme word group affix synonym collocation semantic field