Skip to main navigation Skip to search Skip to main content

Using the Structure of HTML Documents to Improve Retrieval

  • State University of New York Binghamton University

Research output: Contribution to conferencePaperpeer-review

54 Scopus citations

Abstract

The World Wide Web (WWW) is a gigantic information resource, which is growing daily. As more and more data are added to the WWW, it is becoming increasingly difficult to effectively locate useful information from this environment. In this paper, we propose a method for making use of the structures and hyperlinks of HTML documents to improve the effectiveness of retrieving HTML documents. Our study assigns the occurrences of terms in a document collection into six classes according to the tags in which a particular term appears (such as Title, H1-H6, and Anchor). Based on the assignment, we extend the weighting schemes in traditional information retrieval by incorporating different importance factors to terms in different classes. The rationale is that terms appearing in different places of a document may have different significance in identifying the document. For this research we have built a Web based search tool, Webor, created a testbed, and conducted extensive experiments to determine an optimal class importance factor combination. Our study indicates that substantial improvement of retrieval effectiveness can be achieved using this technique.

Original languageEnglish
StatePublished - 1997
Event1st USENIX Symposium on Internet Technologies and Systems, USITS 1997 - Monterey, United States
Duration: Dec 8 1997Dec 11 1997

Conference

Conference1st USENIX Symposium on Internet Technologies and Systems, USITS 1997
Country/TerritoryUnited States
CityMonterey
Period12/8/9712/11/97

Fingerprint

Dive into the research topics of 'Using the Structure of HTML Documents to Improve Retrieval'. Together they form a unique fingerprint.

Cite this