Skip to main navigation Skip to search Skip to main content

A diagnostic evaluation approach for English to Hindi MT using linguistic checkpoints and error rates

  • Renu Balyan
  • , Sudip Kumar Naskar
  • , Antonio Toral
  • , Niladri Chatterjee
  • Dublin City University
  • Indian Institute of Technology Delhi

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

This paper addresses diagnostic evaluation of machine translation (MT) systems for Indian languages, English to Hindi MT in particular, assessing the performance of MT systems on relevant linguistic phenomena (checkpoints). We use the diagnostic evaluation tool DELiC4MT to analyze the performance of MT systems on various PoS categories (e.g. nouns, verbs). The current system supports only word level checkpoints which might not be as helpful in evaluating the translation quality as compared to using checkpoints at phrase level and checkpoints that deal with named entities (NE), inflections, word order, etc. We therefore suggest phrase level checkpoints and NEs as additional checkpoints for DELiC4MT. We further use Hjerson to evaluate checkpoints based on word order and inflections that are relevant for evaluation of MT with Hindi as the target language. The experiments conducted using Hjerson generate overall (document level) error counts and error rates for five error classes (inflectional errors, reordering errors, missing words, extra words, and lexical errors) to take into account the evaluation based on word order and inflections. The effectiveness of the approaches was tested on five English to Hindi MT systems.

Original languageEnglish
Title of host publicationComputational Linguistics and Intelligent Text Processing - 14th International Conference, CICLing 2013, Proceedings
Pages285-296
Number of pages12
EditionPART 2
DOIs
StatePublished - 2013
Event14th Annual Conference on Intelligent Text Processing and Computational Linguistics, CICLing 2013 - Samos, Greece
Duration: Mar 24 2013Mar 30 2013

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
NumberPART 2
Volume7817 LNCS

Conference

Conference14th Annual Conference on Intelligent Text Processing and Computational Linguistics, CICLing 2013
Country/TerritoryGreece
CitySamos
Period03/24/1303/30/13

Keywords

  • DELiC4MT
  • Hjerson
  • automatic evaluation metrics
  • checkpoints
  • diagnostic evaluation
  • errors

Fingerprint

Dive into the research topics of 'A diagnostic evaluation approach for English to Hindi MT using linguistic checkpoints and error rates'. Together they form a unique fingerprint.

Cite this