Skip to main navigation Skip to search Skip to main content

Model based table cell detection and content extraction from degraded document images

  • SUNY Buffalo

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

3 Scopus citations

Abstract

This paper describes a novel method for detection and extraction of contents of table cells from handwritten document images. Given a model of the table and a document image containing a table, the hand-drawn or pre-printed table is detected and the contents of the table cells are extracted automatically. The algorithms described are designed to handle degraded binary document images. The target images may include a wide variety of noise, ranging from clutter noise, salt-and-pepper noise to non-text objects such as graphics and logos. The presented algorithm effectively eliminates extraneous noise and identifies potentially matching table layout candidates by detecting horizontal and vertical table line candidates. A table is represented as a matrix based on the locations of intersections of horizontal and vertical table lines, and a matching algorithm searches for the best table structure that matches the given layout model and using the matching score to eliminate spurious table line candidates. The optimally matched table candidate is then used for cell content extraction. This method was tested on a set of document page images containing tables from the challenge set of the DARPA MADCAT Arabic handwritten document image data. Preliminary results indicate that the method is effective and is capable of reliably extracting text from the table cells.

Original languageEnglish
Title of host publicationProceedings of the Workshop on Document Analysis and Recognition, DAR 2012
Pages62-67
Number of pages6
DOIs
StatePublished - 2012
EventWorkshop on Document Analysis and Recognition, DAR 2012 - Mumbai, India
Duration: Dec 16 2012Dec 16 2012

Publication series

NameACM International Conference Proceeding Series

Conference

ConferenceWorkshop on Document Analysis and Recognition, DAR 2012
Country/TerritoryIndia
CityMumbai
Period12/16/1212/16/12

Keywords

  • handwritten Arabic documents
  • table cell extraction
  • table detection

Fingerprint

Dive into the research topics of 'Model based table cell detection and content extraction from degraded document images'. Together they form a unique fingerprint.

Cite this