Skip to main navigation Skip to search Skip to main content

Training Dataset Curation by L1-Norm Principal-Component Analysis for Support Vector Machines

  • Shruti Shukla
  • , Dimitris A. Pados
  • , George Sklivanitis
  • , Elizabeth Serena Bentley
  • , Michael J. Medley
  • Florida Atlantic University
  • Air Force Research Laboratory

Research output: Contribution to journalArticlepeer-review

Abstract

Support vector machines (SVMs) have been the learning model of choice in numerous classification applications. While SVMs are widely successful in real-world deployments, they remain susceptible to mislabeled examples in training datasets where the presence of few faults can severely affect decision boundaries, thereby affecting the model’s performance on unseen data. In this brief, we develop and describe in implementation detail a novel method based on L1-norm principal-component data analysis and geometry that aims to filter out atypical data instances on a class-by-class basis before the training phase of SVMs and thus provide the classifier with robust support-vector candidates for making classification boundaries. The proposed dataset curation method is entirely data-driven (touch-free), unsupervised, and computationally efficient. Extensive experimental studies on real datasets included in this brief illustrate the L1-norm curation method and demonstrate its efficacy in protecting SVM models from data faults during learning.

Original languageEnglish
Pages (from-to)17499-17507
Number of pages9
JournalIEEE Transactions on Neural Networks and Learning Systems
Volume36
Issue number9
DOIs
StatePublished - 2025

Keywords

  • L-norm
  • dataset curation
  • faulty data
  • mislabeled data
  • outlier resistance
  • principal-component analysis (PCA)
  • rank selection
  • support vector machines (SVMs)

Fingerprint

Dive into the research topics of 'Training Dataset Curation by L1-Norm Principal-Component Analysis for Support Vector Machines'. Together they form a unique fingerprint.

Cite this