Skip to main navigation Skip to search Skip to main content

A vector space model for subjectivity classification in urdu aided by co-training

  • SUNY Buffalo

Research output: Contribution to conferencePaperpeer-review

19 Scopus citations

Abstract

The goal of this work is to produce a classifier that can distinguish subjective sentences from objective sentences for the Urdu language. The amount of labeled data required for training automatic classifiers can be highly imbalanced especially in the multilingual paradigm as generating annotations is an expensive task. In this work, we propose a cotraining approach for subjectivity analysis in the Urdu language that augments the positive set (subjective set) and generates a negative set (objective set) devoid of all samples close to the positive ones. Using the data set thus generated for training, we conduct experiments based on SVM and VSM algorithms, and show that our modified VSM based approach works remarkably well as a sentence level subjectivity classifier.

Original languageEnglish
Pages860-868
Number of pages9
StatePublished - 2010
Event23rd International Conference on Computational Linguistics, Coling 2010 - Beijing, China
Duration: Aug 23 2010Aug 27 2010

Conference

Conference23rd International Conference on Computational Linguistics, Coling 2010
Country/TerritoryChina
CityBeijing
Period08/23/1008/27/10

Fingerprint

Dive into the research topics of 'A vector space model for subjectivity classification in urdu aided by co-training'. Together they form a unique fingerprint.

Cite this