Abstract
Despite several decades of research in document analysis, recognition of unconstrained handwritten documents is still considered a challenging task. Previous research in this area has shown that word recognizers perform adequately on constrained handwritten documents which typically use a restricted vocabulary (lexicon). But in the case of unconstrained handwritten documents, state-of-the-art word recognition accuracy is still below the acceptable limits. The objective of this research is to improve word recognition accuracy on unconstrained handwritten documents by applying a post-processing or OCR correction technique to the word recognition output. In this paper, we present two different methods for this purpose. First, we describe a lexicon reduction-based method by topic categorization of handwritten documents which is used to generate smaller topic-specific lexicons for improving the recognition accuracy. Second, we describe a method which uses topic-specific language models and a maximum-entropy based topic categorization model to refine the recognition output. We present the relative merits of each of these methods and report results on the publicly available IAM database.
| Original language | English |
|---|---|
| Pages (from-to) | 153-164 |
| Number of pages | 12 |
| Journal | International Journal on Document Analysis and Recognition |
| Volume | 12 |
| Issue number | 3 |
| DOIs | |
| State | Published - 2009 |
Keywords
- Document categorization
- Handwritten documents
- Language models
- Lexicon reduction
- OCR correction
- Topic models
- Unconstrained handwriting
Fingerprint
Dive into the research topics of 'Using topic models for OCR correction'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver