TY - GEN
T1 - Cyberbullying detection on instagram with optimal online feature selection
AU - Yao, Mengfan
AU - Chelmis, Charalampos
AU - Zois, Daphney Stavroula
N1 - Publisher Copyright: © 2018 IEEE.
PY - 2018/10/24
Y1 - 2018/10/24
N2 - Cyberbullying has emerged as a large-scale societal problem that demands accurate methods for its detection in an effort to mitigate its detrimental consequences. While automated, data-driven techniques for analyzing and detecting cyberbullying incidents have been developed, the scalability of existing approaches has largely been ignored. At the same time, the complexities underlying cyberbullying behavior (e.g., social context and changing language) make the automatic identification of 'the best subset of features' to use challenging. We address this gap by formulating cyberbullying detection as a sequential hypothesis testing problem. Based on this formulation, we propose a novel algorithm to drastically reduce the number of features used in classification. We demonstrate the utility, scalability and responsiveness of our approach using a real-world dataset from Instagram, the online social media platform with the highest percentage of users reporting experiencing cyberbullying. Our approach improves recall by a staggering 700%, while at the same time reducing the average number of features by up to 99.82% compared to state-of-the-art supervised cyberbullying detection methods, learning approaches that require weak supervision, and traditional offline feature selection and dimensionality reduction techniques.
AB - Cyberbullying has emerged as a large-scale societal problem that demands accurate methods for its detection in an effort to mitigate its detrimental consequences. While automated, data-driven techniques for analyzing and detecting cyberbullying incidents have been developed, the scalability of existing approaches has largely been ignored. At the same time, the complexities underlying cyberbullying behavior (e.g., social context and changing language) make the automatic identification of 'the best subset of features' to use challenging. We address this gap by formulating cyberbullying detection as a sequential hypothesis testing problem. Based on this formulation, we propose a novel algorithm to drastically reduce the number of features used in classification. We demonstrate the utility, scalability and responsiveness of our approach using a real-world dataset from Instagram, the online social media platform with the highest percentage of users reporting experiencing cyberbullying. Our approach improves recall by a staggering 700%, while at the same time reducing the average number of features by up to 99.82% compared to state-of-the-art supervised cyberbullying detection methods, learning approaches that require weak supervision, and traditional offline feature selection and dimensionality reduction techniques.
KW - classification
KW - cyberharassment
KW - online social media
KW - optimization algorithm
KW - selection process
UR - https://www.scopus.com/pages/publications/85057325859
U2 - 10.1109/ASONAM.2018.8508329
DO - 10.1109/ASONAM.2018.8508329
M3 - Conference contribution
T3 - Proceedings of the 2018 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining, ASONAM 2018
SP - 401
EP - 408
BT - Proceedings of the 2018 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining, ASONAM 2018
A2 - Tagarelli, Andrea
A2 - Reddy, Chandan
A2 - Brandes, Ulrik
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 10th IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining, ASONAM 2018
Y2 - 28 August 2018 through 31 August 2018
ER -