Skip to main navigation Skip to search Skip to main content

Document image analysis: A primer

  • Pennsylvania State University
  • Avaya Inc

Research output: Contribution to journalArticlepeer-review

101 Scopus citations

Abstract

Document image analysis refers to algorithms and techniques that are applied to images of documents to obtain a computer-readable description from pixel data. A well-known document image analysis product is the Optical Character Recognition (OCR) software that recognizes characters in a scanned document. OCR makes it possible for the user to edit or search the document's contents. In this paper we briefly describe various components of a document analysis system. Many of these basic building blocks are found in most document analysis systems, irrespective of the particular domain or language to which they are applied. We hope that this paper will help the reader by providing the background necessary to understand the detailed descriptions of specific techniques presented in other papers in this issue.

Original languageEnglish
Pages (from-to)3-22
Number of pages20
JournalSadhana - Academy Proceedings in Engineering Sciences
Volume27
Issue numberPART 1
DOIs
StatePublished - 2002

Keywords

  • Character recognition
  • Document processing
  • Feature analysis
  • Graphics recognition
  • Layout analysis
  • OCR

Fingerprint

Dive into the research topics of 'Document image analysis: A primer'. Together they form a unique fingerprint.

Cite this