Publication Details
Issue: Vol 2, No 9 (2025)
Pages: 144-152
ISSN: 2997-3953

Abstract

This paper examines the fundamental criteria for selecting and analyzing audio materials in the context of corpus linguistics and speech technology, with particular reference to low-resource languages such as Uzbek. Audio corpora play a crucial role in advancing Automatic Speech Recognition (ASR), Natural Language Processing (NLP), and linguistic research, yet their effectiveness depends on the quality and representativeness of the data they contain. Drawing on international best practices and a case study of Kun.uz audio news materials, five key criteria are identified and evaluated: authenticity and naturalness of speech; technical quality and recording standards; sociolinguistic representativeness, including demographic and dialectal diversity; multilayer annotation and metadata enrichment; and ethical and legal responsibility in data collection and sharing. The results demonstrate that while available Uzbek audio resources provide valuable input for ASR development, they remain limited in thematic scope and regional coverage, which restricts their broader applicability. By comparing these findings with established corpora such as TIMIT, SWITCHBOARD, LibriSpeech, and Common Voice, the study highlights both strengths and shortcomings in current practices. The paper concludes with recommendations for expanding thematic domains, improving annotation depth, and ensuring ethical standards, thereby contributing to the creation of more robust, inclusive, and sustainable audio corpora for linguistic and technological innovation.

Keywords
audio corpora speech technology automatic speech recognition uzbek language corpus linguistics data selection criteria