Content-based SMS Classification: Statistical Analysis for the Relationship between Number of Features and Classification Performance



Document title: Content-based SMS Classification: Statistical Analysis for the Relationship between Number of Features and Classification Performance
Journal: Computación y sistemas
Database: PERIÓDICA
System number: 000423290
ISSN: 1405-5546
Authors: 1
1
Institutions: 1Tun Hussein Onn University of Malaysia, Batu Pahat, Johor. Malasia
2Hodeidah University, Computer Science Department, Alduraihimi, Hodeidah. Yemen
Year:
Season: Oct-Dic
Volumen: 21
Number: 4
Country: México
Language: Inglés
Document type: Artículo
Approach: Aplicado, descriptivo
English abstract High dimensionality of the feature space is one of the difficulty that affect short message service (SMS) classification performance. Some studies used feature selection methods to pick up some features, while other studies used the full extracted features. In this work, we aim to analyse the relationship between features size and classification performance. For that, a classification performance comparison was carried out between ten features sizes selected by varies feature selection methods. The used methods were chi-square, Gini index and information gain (IG). Support vector machine was used as a classifier. Area Under the ROC (Receiver Operating Characteristics) Curve between true positive rate and false positive rate was used to measure the classification performance. We used the repeated measures ANOVA at p < 0.05 level to analyse the performance. Experimental results showed that IG method outperformed the other methods in all features sizes. The best result was with 50% of the extracted features. Furthermore, the results explicitly showed that using larger features size in the classification does not mean superior performance but sometimes leads to less classification performance. Therefore, feature selection step should be used. By reducing the used features for the classification, without degrading the classification performance, it means reducing memory usage and classification time
Disciplines: Ciencias de la computación,
Literatura y lingüística
Keyword: Lingüística aplicada,
Clasificación de textos,
Spam,
Filtros digitales,
Selección de características,
Máquinas de soporte vectorial
Keyword: Applied linguistics,
Text classification,
Spam,
Digital filters,
Feature selection,
Support vector machines
Full text: Texto completo (Ver HTML) Texto completo (Ver PDF)