Content Extraction based on Hierarchical Relations in DOM Structures

López, Sergio; Silva, Josep; Insa, David


Document title:	Content Extraction based on Hierarchical Relations in DOM Structures
Journal:	Polibits
Database:	PERIÓDICA
System number:	000355832
ISSN:	1870-9044
Authors:	López, Sergio¹ Silva, Josep¹ Insa, David¹
Institutions:	¹Universitat Politècnica de València, Departamento de Sistemas Informáticos y Computación, Valencia. España
Year:	2012
Number:	45
Country:	México
Language:	Inglés
Document type:	Artículo
Approach:	Analítico, descriptivo
English abstract	This article introduces a new approach for content extraction that exploits the hierarchical inter-relations of the elements in a webpage. Content extraction is a technique used to extract from a webpage the main textual content. This is useful in order to filter out the advertisements and all the additional information that is not part of the main content. The main idea behind our approach is to use the DOM tree as an explicit representation of the inter-relations of the elements in a webpage. Using the information contained in the DOM tree we can identify blocks of content and we can easily determine what of the blocks contains more text. Thanks to this information, the technique achieves a considerable recall and precision. Using the DOM structure for content extraction gives us the benefits of other approaches based on the syntax of the webpage (such as characters, words and tags), but it also gives us a very precise information regarding the related components in a block, thus, producing very cohesive blocks
Disciplines:	Ciencias de la computación
Keyword:	Procesamiento de datos, Análisis y sistematización de la información, Extracción de contenidos, Páginas web, Detección de bloques
Keyword:	Computer science, Data processing, Information analysis, Contents extraction, Web pages, Block detection
Full text:	Texto completo (Ver HTML)

Content Extraction based on Hierarchical Relations in DOM Structures

Wait a moment...