TY - GEN
T1 - Automatic Extraction of Dynamic Record Sections From Search Engine Result Pages
AU - Zhao, Hongkun
AU - Meng, Weiyi
AU - Yu, Clement
N1 - Publisher Copyright: © 2006 Association for Computing Machinery. All rights reserved.
PY - 2006
Y1 - 2006
N2 - A search engine returned result page may contain search results that are organized into multiple dynamically generated sections in response to a user query. Furthermore, such a result page often also contains information irrelevant to the query, such as information related to the hosting site of the search engine. In this paper, we present a method to automatically generate wrappers for extracting search result records from all dynamic sections on result pages returned by search engines. This method has the following novel features: (1) it aims to explicitly identify all dynamic sections, including those that are not seen on sample result pages used to generate the wrapper, and (2) it addresses the issue of correctly differentiating sections and records. Experimental results indicate that this method is very promising. Automatic search result record extraction is critical for applications that need to interact with search engines such as automatic construction and maintenance of metasearch engines and deep Web crawling.
AB - A search engine returned result page may contain search results that are organized into multiple dynamically generated sections in response to a user query. Furthermore, such a result page often also contains information irrelevant to the query, such as information related to the hosting site of the search engine. In this paper, we present a method to automatically generate wrappers for extracting search result records from all dynamic sections on result pages returned by search engines. This method has the following novel features: (1) it aims to explicitly identify all dynamic sections, including those that are not seen on sample result pages used to generate the wrapper, and (2) it addresses the issue of correctly differentiating sections and records. Experimental results indicate that this method is very promising. Automatic search result record extraction is critical for applications that need to interact with search engines such as automatic construction and maintenance of metasearch engines and deep Web crawling.
UR - https://www.scopus.com/pages/publications/85044217577
M3 - Conference contribution
SN - 1595933859
SN - 9781595933850
T3 - VLDB 2006 - Proceedings of the 32nd International Conference on Very Large Data Bases
SP - 989
EP - 1000
BT - VLDB 2006 - Proceedings of the 32nd International Conference on Very Large Data Bases
PB - Association for Computing Machinery
T2 - 32nd International Conference on Very Large Data Bases, VLDB 2006
Y2 - 12 September 2006 through 15 September 2006
ER -