Abstract
Support vector machines (SVMs) have been the learning model of choice in numerous classification applications. While SVMs are widely successful in real-world deployments, they remain susceptible to mislabeled examples in training datasets where the presence of few faults can severely affect decision boundaries, thereby affecting the model’s performance on unseen data. In this brief, we develop and describe in implementation detail a novel method based on L1-norm principal-component data analysis and geometry that aims to filter out atypical data instances on a class-by-class basis before the training phase of SVMs and thus provide the classifier with robust support-vector candidates for making classification boundaries. The proposed dataset curation method is entirely data-driven (touch-free), unsupervised, and computationally efficient. Extensive experimental studies on real datasets included in this brief illustrate the L1-norm curation method and demonstrate its efficacy in protecting SVM models from data faults during learning.
| Original language | English |
|---|---|
| Pages (from-to) | 17499-17507 |
| Number of pages | 9 |
| Journal | IEEE Transactions on Neural Networks and Learning Systems |
| Volume | 36 |
| Issue number | 9 |
| DOIs | |
| State | Published - 2025 |
Keywords
- L-norm
- dataset curation
- faulty data
- mislabeled data
- outlier resistance
- principal-component analysis (PCA)
- rank selection
- support vector machines (SVMs)
Fingerprint
Dive into the research topics of 'Training Dataset Curation by L1-Norm Principal-Component Analysis for Support Vector Machines'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver