Abstract
Foreign language materials on the web are growing at a faster rate than English language materials and it has been predicted that by the end of 1999 the amount of non-English resources on the internet will exceed English resources. A significant portion of the non-English material is in the form of images. We believe that it would be of value if we were to be able to identify the language/script or identify the text by processing the images through English character OCRs that are trained on the English alphabet. The output of the recognizer is analyzed to see if there are unique signature patterns that would help map the resultant English characters or strings to foreign language/script characters or strings. If the language or script used in the image can be identified based on the results of an English recognition system, then the OCR for that script, if available, can be applied and an appropriate Machine Translation Engine can be used to recognize the text. An initial attempt to identify characteristic signature patterns while processing machine-printed Devanagari characters through English OCRs is detailed in this paper.
| Original language | English |
|---|---|
| Pages (from-to) | 305-312 |
| Number of pages | 8 |
| Journal | Proceedings of SPIE - The International Society for Optical Engineering |
| Volume | 3964 |
| State | Published - 2000 |
| Event | Proceedings of the 2000 Internet Imaging - San Jose, CA, USA Duration: Jan 26 2000 → Jan 28 2000 |
Fingerprint
Dive into the research topics of 'Translingual OCR by template correlations'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver