Abstract
The goal of this work is to produce a classifier that can distinguish subjective sentences from objective sentences for the Urdu language. The amount of labeled data required for training automatic classifiers can be highly imbalanced especially in the multilingual paradigm as generating annotations is an expensive task. In this work, we propose a cotraining approach for subjectivity analysis in the Urdu language that augments the positive set (subjective set) and generates a negative set (objective set) devoid of all samples close to the positive ones. Using the data set thus generated for training, we conduct experiments based on SVM and VSM algorithms, and show that our modified VSM based approach works remarkably well as a sentence level subjectivity classifier.
| Original language | English |
|---|---|
| Pages | 860-868 |
| Number of pages | 9 |
| State | Published - 2010 |
| Event | 23rd International Conference on Computational Linguistics, Coling 2010 - Beijing, China Duration: Aug 23 2010 → Aug 27 2010 |
Conference
| Conference | 23rd International Conference on Computational Linguistics, Coling 2010 |
|---|---|
| Country/Territory | China |
| City | Beijing |
| Period | 08/23/10 → 08/27/10 |
Fingerprint
Dive into the research topics of 'A vector space model for subjectivity classification in urdu aided by co-training'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver