TY - GEN
T1 - Detection of fraudulent claims using hierarchical cluster analysis
AU - Khurjekar, Neel
AU - Chou, Chun An
AU - Khasawneh, Mohammad T.
PY - 2015
Y1 - 2015
N2 - The U.S. healthcare system is being affected by fraudulent and abusive activities leading to huge financial expenditure annually. Data mining and predictive analytics-based techniques can help payers validate claims submitted by providers to avoid any suspicious activity. Therefore, this paper suggests a data mining method for detecting fraudulent claims using the Medicare Physician dataset. The data being used for the study is from the Provider Utilization and Payment Data Physician and Other Supplier Public Use File (PUF) dataset prepared by the Centers for Medicare and Medicaid Services (CMS). The dataset contains information on various medical services along with the procedures that have been rendered by Medicare beneficiaries. The provider claims are categorized based on the type and the various services provided by the provider. In this study, a two-step approach is proposed. First, this approach involves the use of multivariate analysis, followed by a cluster analysis to identify the claims that are highly deviating from the group of claims pertaining to the provider type. Using residual analysis, claims having an average error of 85.21% as a result of the first step were identified. In the next step (cluster analysis), fraudulent observations were detected based on an average distance to cluster speed of more than 18,200. The results of both the steps involved in the approach led to the creation of a dataset with suspicious claims that have been submitted for reimbursement by physicians and other care providers that should receive further manual investigation.
AB - The U.S. healthcare system is being affected by fraudulent and abusive activities leading to huge financial expenditure annually. Data mining and predictive analytics-based techniques can help payers validate claims submitted by providers to avoid any suspicious activity. Therefore, this paper suggests a data mining method for detecting fraudulent claims using the Medicare Physician dataset. The data being used for the study is from the Provider Utilization and Payment Data Physician and Other Supplier Public Use File (PUF) dataset prepared by the Centers for Medicare and Medicaid Services (CMS). The dataset contains information on various medical services along with the procedures that have been rendered by Medicare beneficiaries. The provider claims are categorized based on the type and the various services provided by the provider. In this study, a two-step approach is proposed. First, this approach involves the use of multivariate analysis, followed by a cluster analysis to identify the claims that are highly deviating from the group of claims pertaining to the provider type. Using residual analysis, claims having an average error of 85.21% as a result of the first step were identified. In the next step (cluster analysis), fraudulent observations were detected based on an average distance to cluster speed of more than 18,200. The results of both the steps involved in the approach led to the creation of a dataset with suspicious claims that have been submitted for reimbursement by physicians and other care providers that should receive further manual investigation.
KW - Cluster analysis
KW - Fraudulent claims
KW - Medicare
KW - Multivariate analysis
KW - Predictive analytics
UR - https://www.scopus.com/pages/publications/84971016770
M3 - Conference contribution
T3 - IIE Annual Conference and Expo 2015
SP - 2388
EP - 2396
BT - IIE Annual Conference and Expo 2015
PB - Institute of Industrial Engineers
T2 - IIE Annual Conference and Expo 2015
Y2 - 30 May 2015 through 2 June 2015
ER -