Combining N-grams and graph convolution for text classification

dc.authoridBakal, Mehmet/0000-0003-2897-3894
dc.contributor.authorSen, Tarik Uveys
dc.contributor.authorYakit, Mehmet Can
dc.contributor.authorGumus, Mehmet Semih
dc.contributor.authorAbar, Orhan
dc.contributor.authorBakal, Gokhan
dc.date.accessioned2025-08-12T08:26:56Z
dc.date.issued2025
dc.departmentOsmaniye Korkut Ata Üniversitesi
dc.description.abstractText classification, a cornerstone of natural language processing (NLP), finds applications in diverse areas, from sentiment analysis to topic categorization. While deep learning models have recently dominated the field, traditional n-gram-driven approaches often struggle to achieve comparable performance, particularly on large datasets. This gap largely stems from deep learning' s superior ability to capture contextual information through word embeddings. This paper explores a novel approach to leverage the often-overlooked power of n-gram features for enriching word representations and boosting text classification accuracy. We propose a method that transforms textual data into graph structures, utilizing discriminative n-gram series to establish long-range relationships between words. By training a graph convolution network on these graphs, we derive contextually enhanced word embeddings that encapsulate dependencies extending beyond local contexts. Our experiments demonstrate that integrating these enriched embeddings into an long-short term memory (LSTM) model for text classification leads to around 2% improvements in classification performance across diverse datasets. This achievement highlights the synergy of combining traditional n-gram features with graph-based deep learning techniques for building more powerful text classifiers.
dc.description.sponsorshipScientific and Technological Research Council of Turkiye (TUBITAK) [122E103]
dc.description.sponsorshipThis research is funded by the Scientific and Technological Research Council of Turkiye (TUBITAK) through 3501 Career Development Program with grant number 122E103. The authors also express their gratitude to Google Cloud Services for providing academic credit support that facilitated portions of this work.
dc.identifier.doi10.1016/j.asoc.2025.113092
dc.identifier.issn1568-4946
dc.identifier.issn1872-9681
dc.identifier.scopus2-s2.0-105001822010
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.1016/j.asoc.2025.113092
dc.identifier.urihttps://hdl.handle.net/20.500.12502/5215
dc.identifier.volume175
dc.identifier.wosWOS:001464755900001
dc.identifier.wosqualityQ1
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherElsevier
dc.relation.ispartofApplied Soft Computing
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_WOS_20250812
dc.subjectText-graph transformation
dc.subjectGraph convolution network
dc.subjectDeep learning
dc.subjectText mining
dc.subjectGraph mining
dc.titleCombining N-grams and graph convolution for text classification
dc.typeArticle

Dosyalar