Evaluation of Feature Extraction Techniques in Automatic Authorship Attribution



Document title: Evaluation of Feature Extraction Techniques in Automatic Authorship Attribution
Journal: Computación y sistemas
Database:
System number: 000560803
ISSN: 1405-5546
Authors: 1
2
2
1
3
1
Institutions: 1Tecnológico Nacional de México, México
2Instituto Politécnico Nacional, Escuela Superior de Ingeniería Mecánica y Eléctrica, México
3Instituto Tecnológico Superior de los Ríos, Balancán, Tabasco. México
Year:
Season: Abr-Jun
Volumen: 27
Number: 2
Pages: 371-377
Country: México
Language: Inglés
English abstract There are two main approaches to automatic text classification: content-based classification and style-based classification. With content-based text classification, the topic of a document (politics, sports, health) or fake news is detected. On the other hand, Style-based text classification is used to detect the gender or age of an author, author identification, and authorship attribution. In style-based classification, the set of words defines the author’s vocabulary, which contains several hundred words. In this work, the words are known as dimensions. Texts generate high-dimensional vectors. Multiple works have shown that a large number of dimensions decreases the performance of classifiers. To reduce dimensions there are selection and extraction techniques. This article discusses the use of extraction techniques, which create low-dimensional vectors from combinations of the high-dimensional vector. Due to the development of Deep Learning networks, the use of dimensión reduction techniques has decreased because these networks perform dimensión reduction automatically. However, in Machine Learning such techniques are still used intensively. Motivated by the above, in this paper, the Principal Component Analysis (PCA) and Latent Semantic Analysis (LSA) dimensión reduction algorithms are proposed for the identification of texts written by 14 authors of the Corpus PAN 2012. The texts were divided into sequences of 10, 20, and 30 words called sentences. Likewise, blocks of texts made up of 100 sentences were created. The supervised classification was performed with the Nearest Neighbors (KNN), Support Vector Machines (SVM) and Logistic Regression (LR) algorithms using the accuracy metric. The results showed that the reduction of dimensions with PCA and the LR and SVM classifiers achieved better results than other similar works of the state of the art using the same corpus.
Keyword: Dimension reduction,
Feature extraction,
Authorship attribution,
Machine learning
Full text: Texto completo (Ver HTML) Texto completo (Ver PDF)