Author Verification Using a Semantic Space Model



Título del documento: Author Verification Using a Semantic Space Model
Revue: Computación y Sistemas
Base de datos: PERIÓDICA
Número de sistema: 000423242
ISSN: 1405-5546
Autores: 1
1
Instituciones: 1Instituto Politécnico Nacional, Centro de Investigación en Computación, Ciudad de México. México
2Tecnológico de Estudios Superiores de Tianguistenco, Santiago Tianguistenco, Estado de México. México
Año:
Periodo: Abr-Jun
Volumen: 21
Número: 2
País: México
Idioma: Inglés
Tipo de documento: Artículo
Enfoque: Aplicado, descriptivo
Resumen en inglés In this work we propose to solve the author verification problem using a semantic space model through Latent Dirichlet Allocation (LDA). We experiment with the corpus used in the author identification tasks at PAN 2014 and PAN 2015. These datasets consist of subsets in the following languages: English, Spanish, Dutch and Greek. Each problem contained in these corpora is formed by one to five known documents which were written by one author and one unknown document. The task is to predict whether the unknown document was written by the author who wrote the known documents. We processed the documents in the dataset and captured the fingerprint of authors by generating a probabilistic distribution of words in the documents. In PAN 2015 classification, we achieved 81.6%, 75.4%, 74.1%, 67.1% accuracy for each English, Spanish, Dutch and Greek subset respectively. In particular for the English subset, we outreached the best result reported in both competitions
Disciplinas: Literatura y lingüística,
Bibliotecología y ciencia de la información
Palabras clave: Tecnología de la información,
Lingüística aplicada,
Verificación de autoría,
Modelo de espacio semántico
Keyword: Information technology,
Applied linguistics,
Authorship verification,
Semantic space model
Texte intégral: Texto completo (Ver HTML) Texto completo (Ver PDF)