SciELO - Scientific Electronic Library Online

 
 número25Algoritmo de predicción del consumo de combustible para mezcla de etanol anhídrido en ciudades de alturaAnálisis de la eficiencia de un disco de freno convencional ventilado con respecto a un disco hiperventilado mediante mecanizado índice de autoresíndice de assuntospesquisa de artigos
Home Pagelista alfabética de periódicos  

Serviços Personalizados

Journal

Artigo

Indicadores

Links relacionados

  • Não possue artigos similaresSimilares em SciELO

Compartilhar


Ingenius. Revista de Ciencia y Tecnología

versão On-line ISSN 1390-860Xversão impressa ISSN 1390-650X

Resumo

ESCALONA ESCALONA, Yosveni. Algorithms for Table Structure Recognition. Ingenius [online]. 2021, n.25, pp.50-61. ISSN 1390-860X.  https://doi.org/10.17163/ings.n25.2021.05.

Tables are widely adopted to organize and publish data. For example, the Web has an enormous number of tables, published in HTML, embedded in PDF documents, or that can be simply downloaded from Web pages. However, tables are not always easy to interpret due to the variety of features and formats used. Indeed, a large number of methods and tools have been developed to interpreted tables. This work presents the implementation of an algorithm, based on Conditional Random Fields (CRFs), to classify the rows of a table as header rows, data rows or metadata rows. The implementation is complemented by two algorithms for table recognition in a spreadsheet document, respectively based on rules and on region detection. Finally, the work describes the results and the benefits obtained by applying the implemented algorithm to HTML tables, obtained from the Web, and to spreadsheet tables, downloaded from the Brazilian National Petroleum Agency.

Palavras-chave : Tabular Data; HTML Tables; Spreadsheets; Conditional Random Fields; Machine Learning; Algorithm.

        · resumo em Espanhol     · texto em Espanhol     · Espanhol ( pdf )