Hierarchical vs. flat n-gram-based text categorization: can we do better?

Jelena Graovac; Jovana Kovačević; Gordana Pavlović-Lažetić

Jelena Graovac ; Jovana Kovačević ; Gordana Pavlović-Lažetić

Computer Science and Information Systems, Tome 14 (2017) no. 1

Cet article a éte moissonné depuis la source Computer Science and Information Systems website

Voir la notice de l'article

Résumé

Hierarchical text categorization (HTC) refers to assigning a text document to one or more most suitable categories from a hierarchical category space. In this paper we present two HTC techniques based on kNN and SVM machine learning techniques for categorization process and byte n-gram based document representation. They are fully language independent and do not require any text preprocessing steps, or any prior information about document content or language. The effectiveness of the presented techniques and their language independence are demonstrated in experiments performed on five tree-structured benchmark category hierarchies that differ in many aspects: Reuters-Hier1, Reuters-Hier2, 15NGHier and 20NGHier in English and TanCorpHier in Chinese. The results obtained are compared with the corresponding flat categorization techniques applied to leaf level categories of the considered hierarchies. While kNN-based flat text categorization produced slightly better results than kNN-based HTC on the largest TanCorpHier and 20NGHier datasets, SVM-based HTC results do not considerably differ from the corresponding flat techniques, due to shallow hierarchies; still, they outperform both kNN-based flat and hierarchical categorization on all corpora except the smallest Reuters-Hier1 and Reuters-Hier2 datasets. Formal evaluation confirmed that the proposed techniques obtained state-of-the-art results.

Keywords: hierarchical text categorization, n-grams, kNN, SVM

@article{CSIS_2017_14_1_a6,
     author = {Jelena Graovac and Jovana Kova\v{c}evi\'c and Gordana Pavlovi\'c-La\v{z}eti\'c},
     title = {Hierarchical vs. flat n-gram-based text categorization: can we do better?},
     journal = {Computer Science and Information Systems},
     year = {2017},
     volume = {14},
     number = {1},
     url = {http://geodesic.mathdoc.fr/item/CSIS_2017_14_1_a6/}
}

TY  - JOUR
AU  - Jelena Graovac
AU  - Jovana Kovačević
AU  - Gordana Pavlović-Lažetić
TI  - Hierarchical vs. flat n-gram-based text categorization: can we do better?
JO  - Computer Science and Information Systems
PY  - 2017
VL  - 14
IS  - 1
UR  - http://geodesic.mathdoc.fr/item/CSIS_2017_14_1_a6/
ID  - CSIS_2017_14_1_a6
ER  -

%0 Journal Article
%A Jelena Graovac
%A Jovana Kovačević
%A Gordana Pavlović-Lažetić
%T Hierarchical vs. flat n-gram-based text categorization: can we do better?
%J Computer Science and Information Systems
%D 2017
%V 14
%N 1
%U http://geodesic.mathdoc.fr/item/CSIS_2017_14_1_a6/
%F CSIS_2017_14_1_a6

Jelena Graovac; Jovana Kovačević; Gordana Pavlović-Lažetić. Hierarchical vs. flat n-gram-based text categorization: can we do better?. Computer Science and Information Systems, Tome 14 (2017) no. 1. http://geodesic.mathdoc.fr/item/CSIS_2017_14_1_a6/

Parcourir par

Geodesic

Parcourir par