TY - GEN
T1 - Graph-based text representation and knowledge discovery
AU - Jin, Wei
AU - Srihari, Rohini K.
PY - 2007
Y1 - 2007
N2 - For information retrieval and text-mining, a robust scalable framework is required to represent the information extracted from documents and enable visualization and query of such information. One very widely used model is the vector space model which is based on the bag-of-words approach. However, it suffers from the fact that it loses important information about the original text, such as information about the order of the terms in the text or about the frontiers between sentences or paragraphs. In this paper, we propose a graph-based text representation, which is capable of capturing (i) Term order (ii) Term frequency (iii) Term co-occurrence (iv) Term context in documents. We also apply the graph model into our text mining task, which is to discover unapparent associations between two and more concepts (e.g. individuals) from a large text corpus. Counterterrorism corpus is used to evaluate the performance of various retrieval models, which demonstrates feasibility and effectiveness of graphic text representation in information retrieval and text mining.
AB - For information retrieval and text-mining, a robust scalable framework is required to represent the information extracted from documents and enable visualization and query of such information. One very widely used model is the vector space model which is based on the bag-of-words approach. However, it suffers from the fact that it loses important information about the original text, such as information about the order of the terms in the text or about the frontiers between sentences or paragraphs. In this paper, we propose a graph-based text representation, which is capable of capturing (i) Term order (ii) Term frequency (iii) Term co-occurrence (iv) Term context in documents. We also apply the graph model into our text mining task, which is to discover unapparent associations between two and more concepts (e.g. individuals) from a large text corpus. Counterterrorism corpus is used to evaluate the performance of various retrieval models, which demonstrates feasibility and effectiveness of graphic text representation in information retrieval and text mining.
KW - Information retrieval
KW - Knowledge discovery
KW - Text representation
UR - https://www.scopus.com/pages/publications/35248894286
U2 - 10.1145/1244002.1244182
DO - 10.1145/1244002.1244182
M3 - Conference contribution
SN - 1595934804
SN - 9781595934802
T3 - Proceedings of the ACM Symposium on Applied Computing
SP - 807
EP - 811
BT - Proceedings of the 2007 ACM Symposium on Applied Computing
T2 - 2007 ACM Symposium on Applied Computing
Y2 - 11 March 2007 through 15 March 2007
ER -