Skip to main navigation Skip to search Skip to main content

Polyglot: Distributed word representations for multilingual NLP

  • Stony Brook University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

283 Scopus citations

Abstract

Distributed word representations (word embeddings) have recently contributed to competitive performance in language modeling and several NLP tasks. In this work, we train word embeddings for more than 100 languages using their corresponding Wikipedias. We quantitatively demonstrate the utility of our word embeddings by using them as the sole features for training a part of speech tagger for a subset of these languages. We find their performance to be competitive with near state-of-art methods in English, Danish and Swedish. Moreover, we investigate the semantic features captured by these embeddings through the proximity of word groupings. We will release these embeddings publicly to help researchers in the development and enhancement of multilingual applications.

Original languageEnglish
Title of host publicationCoNLL 2013 - 17th Conference on Computational Natural Language Learning, Proceedings
PublisherAssociation for Computational Linguistics (ACL)
Pages183-192
Number of pages10
ISBN (Electronic)9781937284701
StatePublished - 2013
Event17th Conference on Computational Natural Language Learning, CoNLL 2013 - Sofia, Bulgaria
Duration: Aug 8 2013Aug 9 2013

Publication series

NameCoNLL 2013 - 17th Conference on Computational Natural Language Learning, Proceedings

Conference

Conference17th Conference on Computational Natural Language Learning, CoNLL 2013
Country/TerritoryBulgaria
CitySofia
Period08/8/1308/9/13

Fingerprint

Dive into the research topics of 'Polyglot: Distributed word representations for multilingual NLP'. Together they form a unique fingerprint.

Cite this