Text similarity detection in agglutinative languages: a case study of Kazakh using hybrid n-gram and semantic models

dc.contributor.authorБілощицька Світлана
dc.contributor.authorТлеубаєва Арайлим
dc.contributor.authorКучанський Олександр
dc.contributor.authorБілощицький Андрій
dc.contributor.authorАндрашко Юрій
dc.contributor.authorТоксанов Сапар
dc.contributor.authorМухатаєв Айдос
dc.contributor.authorШаріпова Салтанат
dc.date.accessioned2025-12-07T17:42:11Z
dc.date.issued2025
dc.description.abstractThis study presents an advanced hybrid approach for detecting near-duplicate texts in the Kazakh language, addressing the specific challenges posed by its agglutinative morphology. The proposed method combines statistical and semantic techniques, including N-gram analysis, TF-IDF, LSH, LSA, and LDA, and is benchmarked against the bert-base-multilingual-cased model. Experiments were conducted on the purpose-built Arailym-aitu/KazakhTextDuplicates corpus, which contains over 25,000 manually modified text fragments using typical techniques, such as paraphrasing, word order changes, synonym substitution, and morphological transformations. The results show that the hybrid model achieves a precision of 1.00, a recall of 0.73, and an F1-score of 0.84, significantly outperforming traditional N-gram and TF-IDF approaches and demonstrating comparable accuracy to the BERT model while requiring substantially lower computational resources. The hybrid model proved highly effective in detecting various types of near-duplicate texts, including paraphrased and structurally modified content, making it suitable for practical applications in academic integrity verification, plagiarism detection, and intelligent text analysis. Moreover, this study highlights the potential of lightweight hybrid architectures as a practical alternative to large transformer-based models, particularly for languages with limited annotated corpora and linguistic resources. It lays the foundation for future research in cross-lingual duplicate detection and deep model adaptation for the Kazakh language.
dc.description.sponsorshipThis paper was written in the framework of the state order to implement the research project, IRN No. AP23490123 «Development of a system to detect plagiarism using combined methods, models for finding near-duplicate, focusing on the Kazakh language»
dc.identifier.citationBiloshchytska S., Tleubayeva A., Kuchanskyi O., Biloshchytskyi A., Andrashko Y., Toxanov S., Mukhatayev A., Sharipova S. Text similarity detection in agglutinative languages: a case study of Kazakh using hybrid n-gram and semantic models. Applied Sciences. Vol. 15, Issue 12. 2025. Pub. 6707. DOI: https://doi.org/10.3390/app15126707
dc.identifier.urihttps://dspace.uzhnu.edu.ua/handle/lib/80512
dc.language.isoen
dc.pubTypeСтаття
dc.publisherApplied Sciences
dc.relation.ispartofseries15(12)
dc.subjectanti-plagiarism
dc.subjectKazakh language
dc.subjectcombined models
dc.subjecttext data analysis
dc.subjectnear duplicates
dc.subjectsemantic analysis
dc.subjectacademic integrity
dc.subjectintelligent analysis system
dc.titleText similarity detection in agglutinative languages: a case study of Kazakh using hybrid n-gram and semantic models
dc.typeText

Файли

Контейнер файлів

Зараз показуємо 1 - 1 з 1
Вантажиться...
Ескіз
Назва:
applsci-15-06707 (2).pdf
Розмір:
1.91 MB
Формат:
Adobe Portable Document Format

Ліцензійна угода

Зараз показуємо 1 - 1 з 1
Вантажиться...
Ескіз
Назва:
license.txt
Розмір:
6.33 KB
Формат:
Item-specific license agreed upon to submission
Опис: