TY - GEN
T1 - The performance of sequential and parallel implementations of fp-growth in mining a pharmacy database
AU - Khader, Nourma
AU - Yoon, Sang Won
PY - 2015
Y1 - 2015
N2 - This research studies the performance of a data mining approach through the use of Hadoop and MapReduce paradigm in a novel application. The application involves extracting useful knowledge from an actual pharmacy transactional database to improve different pharmacy management strategies. The study involves the use of a Frequent Pattern Growth (FP-growth) approach, which is one of the algorithms of Association Rule Mining (ARM) to find the association rules among the prescriptions. Since the application of FP-growth on a pharmacy database is novel, sequential and parallel FP-growth are tested. Furthermore, experiments are conducted to examine the effect of null transactions on the performance of the FP-growth algorithm, in which null transactions are single item purchases transactions. The goal is to seek the best implementation that reduces the execution time of FP-growth on such an application. Two datasets are tested: 1) an original transactional dataset that includes 3,828,903 transactions, and 2) a dataset that only includes orders of multiple prescriptions of 725,991 transactions. Results indicate that the performance of the sequential and parallel implementation of FP-growth is dependent on the predetermined minimum support threshold value, ξ. Moreover, excluding the null transactions from the datasets allows for a faster execution of FP-growth.
AB - This research studies the performance of a data mining approach through the use of Hadoop and MapReduce paradigm in a novel application. The application involves extracting useful knowledge from an actual pharmacy transactional database to improve different pharmacy management strategies. The study involves the use of a Frequent Pattern Growth (FP-growth) approach, which is one of the algorithms of Association Rule Mining (ARM) to find the association rules among the prescriptions. Since the application of FP-growth on a pharmacy database is novel, sequential and parallel FP-growth are tested. Furthermore, experiments are conducted to examine the effect of null transactions on the performance of the FP-growth algorithm, in which null transactions are single item purchases transactions. The goal is to seek the best implementation that reduces the execution time of FP-growth on such an application. Two datasets are tested: 1) an original transactional dataset that includes 3,828,903 transactions, and 2) a dataset that only includes orders of multiple prescriptions of 725,991 transactions. Results indicate that the performance of the sequential and parallel implementation of FP-growth is dependent on the predetermined minimum support threshold value, ξ. Moreover, excluding the null transactions from the datasets allows for a faster execution of FP-growth.
KW - Association rules
KW - FP-growth
KW - Hadoop
KW - Null transactions
UR - https://www.scopus.com/pages/publications/84970990752
M3 - Conference contribution
T3 - IIE Annual Conference and Expo 2015
SP - 2601
EP - 2610
BT - IIE Annual Conference and Expo 2015
PB - Institute of Industrial Engineers
T2 - IIE Annual Conference and Expo 2015
Y2 - 30 May 2015 through 2 June 2015
ER -