Revista: | Computación y sistemas |
Base de datos: | |
Número de sistema: | 000560722 |
ISSN: | 1405-5546 |
Autores: | Dömötör, Andrea3 Kákonyi, Tibor1 Yang, Zijian Győző2 |
Instituciones: | 1MTA-PPKE Hungarian Language Technology Research Group, Budapest. Hungría 2Pázmány Péter Catholic University, Faculty of Information Technology and Bionics, Central Hungary. Hungría 3Pázmány Péter Catholic University, Faculty of Humanities and Social Studies, Central Hungary. Hungría |
Año: | 2022 |
Periodo: | Jul-Sep |
Volumen: | 26 |
Número: | 3 |
Paginación: | 1293-1299 |
País: | México |
Idioma: | Inglés |
Tipo de documento: | Artículo |
Resumen en inglés | Genre identification is an important task in natural language processing that can be useful for many practical and research purposes. The challenge of this task is that genre is not a homogeneous and unequivocal property of the texts and it is often hard to separate from the topic. In this paper we compare the performance of two different automatic genre identification methods. We classified six text types: literary, academic, legal, press, spoken and personal. In one part of our research we did experiments with traditional machine learning methods using linguistic, n-gram and error features. In the other part we tested the same task with a word embedding based neural network. In this part we did experiments with different training data (words only, POS-tags only, words and POS-tags etc.). Our results revealed that neural network is a suitable method for this task while traditional machine learning showed significantly lower performance. We gained high (around 70%) accuracy with our word embedding based method. The results of the different text categories seemed to depend on the stylistic properties of the studied genres. |
Disciplinas: | Ciencias de la computación |
Palabras clave: | Inteligencia artificial |
Keyword: | Artificial intelligence |
Texto completo: | Texto completo (Ver HTML) Texto completo (Ver PDF) |