Linguistic Foundations for The Automatic Identification of Prefixal Units in Uzbek
DOI:
https://doi.org/10.55640/eijp-06-06-25Keywords:
Uzbek language, computational morphology, prefix identificationAbstract
The article sets out the linguistic foundations required for the automatic identification of prefixal units in Uzbek texts. Uzbek is an agglutinative language whose computational morphology has been built almost entirely around suffixation, and the preposed formatives anti-, mikro-, makro-, kiber-, eko-, no-, be- and ser- are therefore either segmented incorrectly or not segmented at all by existing analysers. The paper describes four sources of error: the homography of a prefix with the initial syllable of an unrelated root, the homography of a prefix with a free word, the graphic variation between solid, hyphenated and separated spelling, and the coexistence of the Latin and Cyrillic scripts with several competing transliteration conventions. On this basis a layered identification procedure is proposed, combining a closed lexicon of formatives with their combinatorial restrictions, a stop-list of homographic roots, a finite-state morphotactic component and a statistical confidence score derived from the productivity of the formative. The procedure is designed to be compatible with existing finite-state analysers of Turkic languages and with subword segmentation used in neural models.
References
Hojiyev A. O‘zbek tili so‘z yasalishi. – Toshkent: O‘qituvchi, 1989. – 168 b.
G‘ulomov A., Tixonov A., Qo‘ng‘urov R. O‘zbek tili morfem lug‘ati. – Toshkent: O‘qituvchi, 1977. – 464 b.
Jurafsky D., Martin J. H. Speech and Language Processing. 3rd ed. – Stanford: Stanford University, 2023.
Oflazer K. Two-level Description of Turkish Morphology. // Literary and Linguistic Computing. – 1994. – Vol. 9, No. 2. – P. 137–148.
Beesley K. R., Karttunen L. Finite State Morphology. – Stanford: CSLI Publications, 2003. – 611 p.
Sennrich R., Haddow B., Birch A. Neural Machine Translation of Rare Words with Subword Units. // Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics. – Berlin, 2016. – P. 1715–1725.
Creutz M., Lagus K. Unsupervised Models for Morpheme Segmentation and Morphology Learning. // ACM Transactions on Speech and Language Processing. – 2007. – Vol. 4, No. 1. – P. 1–34.
Koskenniemi K. Two-Level Morphology: A General Computational Model for Word-Form Recognition and Production. – Helsinki: University of Helsinki, 1983. – 160 p.
Booij G. The Grammar of Words: An Introduction to Linguistic Morphology. 2nd ed. – Oxford: Oxford University Press, 2007. – 353 p.
Rahmatullayev Sh. Hozirgi adabiy o‘zbek tili. – Toshkent: Universitet, 2006. – 464 b.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Narzullayeva Muborak Sherzod qizi

This work is licensed under a Creative Commons Attribution 4.0 International License.
Individual articles are published Open Access under the Creative Commons Licence: CC-BY 4.0.