A Study on Effective Internet Data Extraction through Layout Detection

  • Sun Bok-Keun (Dept. Computer Engineering, Asan, Hoseo University) ;
  • Han Kwang-Rok (Dept. Computer Engineering, Asan, Hoseo University)
  • Published : 2005.10.01

Abstract

Currently most Internet documents including data are made based on predefined templates, but templates are usually formed only for main data and are not helpful for information retrieval against indexes, advertisements, header data etc. Templates in such forms are not appropriate when Internet documents are used as data for information retrieval. In order to process Internet documents in various areas of information retrieval, it is necessary to detect additional information such as advertisements and page indexes. Thus this study proposes a method of detecting the layout of Web pages by identifying the characteristics and structure of block tags that affect the layout of Web pages and calculating distances between Web pages. This method is purposed to reduce the cost of Web document automatic processing and improve processing efficiency by providing information about the structure of Web pages using templates through applying the method to information retrieval such as data extraction.

Keywords