Skip to main navigation Skip to search Skip to main content

Advanced techniques in web data pre-processing and cleaning

  • Universidad de Chile

Research output: Chapter in Book/Report/Conference proceedingChapterpeer-review

3 Scopus citations

Abstract

Central to successful e-business is the construction of web sites that attract users, capture user preferences, and entice them into making a purchase. Web mining is diverse data mining applied to categorize both the content and structure of web sites with the goal of aiding e-business. Web mining requires knowledge of the web site structure (hyperlink graph), the web content (vector model) and user sessions (the sequence of pages visited by each user to a site). Much of the data for web mining can be noisy. The origin of the noise comes from many sources, for example, undocumented changes to the web site structure and content, a different understanding of the text and media semantic, and web logs without individual user identification. There may not be any record of the number of times a specific page has been visited in a session as page is stored on a proxy or web browser cache. Such noise presents a challenge for web mining. This chapter presents issues with and approaches for cleaning web data in preparation for web mining analysis.

Original languageEnglish
Title of host publicationAdvanced Techniques in Web Intelligence - 1
EditorsJuan Velasquez, Lakhmi Jain
Pages19-48
Number of pages30
DOIs
StatePublished - 2010

Publication series

NameStudies in Computational Intelligence
Volume311

Fingerprint

Dive into the research topics of 'Advanced techniques in web data pre-processing and cleaning'. Together they form a unique fingerprint.

Cite this