English Dataset For Automatic Forum Extraction


Título del documento:	English Dataset For Automatic Forum Extraction
Revista:	Computación y sistemas
Base de datos:
Número de sistema:	000560438
ISSN:	1405-5546
Autors:	Sido, Jakub¹ Konopík, Miloslav¹ Pražák, Ondřej¹
Institucions:	¹University of West Bohemia, Faculty of Applied Sciences, Plzeň. República Checa
Any:	2019
Període:	Jul-Sep
Volum:	23
Número:	3
Paginació:	765-771
País:	México
Idioma:	Inglés
Tipo de documento:	Artículo
Resumen en inglés	This paper describes the process of collecting, maintaining and exploiting an English dataset of web discussions. The dataset consists of many web discussions with hand-annotated posts in the context of a tree structure of a web page. Each post consists of username, date, text, and citations used by its author. The dataset contains 79 different websites with at least 500 pages from each. Each web page consists of a tree structure of HTML tags with texts taken from selected web pages. In the paper, we also describe algorithms trained on the dataset. The algorithms employ basic architectures (such as a bag of words with an SVM classifier and an LSTM network) to set a baseline for the dataset.
Disciplines	Ciencias de la computación
Paraules clau:	Inteligencia artificial
Keyword:	Information retrieval, Web discussion, Artificial intelligence
Text complet:	Texto completo (Ver HTML) Texto completo (Ver PDF)

Esperi un moment...