Skip to main navigation Skip to search Skip to main content

Online AUC optimization for sparse high-dimensional datasets

  • Stony Brook University
  • Stony Brook University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

6 Scopus citations

Abstract

The Area Under the ROC Curve (AUC) is a widely used performance measure for imbalanced classification arising from many application domains where high-dimensional sparse data is abundant. In such cases, each d dimensional sample has only k non-zero features with kll d, and data arrives sequentially in a streaming form. Current online AUC optimization algorithms have high per-iteration cost mathcal{O}(d) and usually produce non-sparse solutions in general, and hence are not suitable for handling the data challenge mentioned above. In this paper, we aim to directly optimize the AUC score for high-dimensional sparse datasets under online learning setting and propose a new algorithm, FTRL-AUC. Our proposed algorithm can process data in an online fashion with a much cheaper per-iteration cost mathcal{O}(k), making it amenable for high-dimensional sparse streaming data analysis. Our new algorithmic design critically depends on a novel reformulation of the U-statistics AUC objective function as the empirical saddle point reformulation, and the innovative introduction of the 'lazy update' rule so that the per-iteration complexity is dramatically reduced from mathcal{O}(d) to mathcal{O}(k). Furthermore, FTRL-AUC can inherently capture sparsity more effectively by applying a generalized Follow-The-Regularized-Leader (FTRL) framework. Experiments on real-world datasets demonstrate that FTRL-AUC significantly improves both run time and model sparsity while achieving competitive AUC scores compared with the state-of-the-art methods. Comparison with the online learning method for logistic loss demonstrates that FTRL-AUC achieves higher AUC scores especially when datasets are imbalanced. Experiments on real-world datasets demonstrate that FTRL-AUC significantly improves both run time and model sparsity while achieving competitive AUC scores compared with the state-of-the-art methods. Comparison with the online learning method for logistic loss demonstrates that FTRL-AUC achieves higher AUC scores especially when datasets are imbalanced.

Original languageEnglish
Title of host publicationProceedings - 20th IEEE International Conference on Data Mining, ICDM 2020
EditorsClaudia Plant, Haixun Wang, Alfredo Cuzzocrea, Carlo Zaniolo, Xindong Wu
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages881-890
Number of pages10
ISBN (Electronic)9781728183169
DOIs
StatePublished - Nov 2020
Event20th IEEE International Conference on Data Mining, ICDM 2020 - Virtual, Online, Italy
Duration: Nov 17 2020Nov 20 2020

Publication series

NameProceedings - IEEE International Conference on Data Mining, ICDM
Volume2020-November

Conference

Conference20th IEEE International Conference on Data Mining, ICDM 2020
Country/TerritoryItaly
CityVirtual, Online
Period11/17/2011/20/20

Keywords

  • AUC optimization
  • Follow-The-Regularized-Leader
  • Online learning
  • Sparsity

Fingerprint

Dive into the research topics of 'Online AUC optimization for sparse high-dimensional datasets'. Together they form a unique fingerprint.

Cite this