A Study on Effective Internet Data Extraction through Layout Detection

Sun Bok-Keun;Han Kwang-Rok;

International Journal of Contents

Volume 1 Issue 2
/
Pages.5-9
/
2005
/
1738-6764(pISSN)
/
2093-7504(eISSN)

The Korea Contents Association (한국콘텐츠학회)

A Study on Effective Internet Data Extraction through Layout Detection

Sun Bok-Keun (Dept. Computer Engineering, Asan, Hoseo University) ;
Han Kwang-Rok (Dept. Computer Engineering, Asan, Hoseo University)

Published : 2005.10.01

PDF

Download PDF

⟨ Previous Next ⟩

Abstract

Currently most Internet documents including data are made based on predefined templates, but templates are usually formed only for main data and are not helpful for information retrieval against indexes, advertisements, header data etc. Templates in such forms are not appropriate when Internet documents are used as data for information retrieval. In order to process Internet documents in various areas of information retrieval, it is necessary to detect additional information such as advertisements and page indexes. Thus this study proposes a method of detecting the layout of Web pages by identifying the characteristics and structure of block tags that affect the layout of Web pages and calculating distances between Web pages. This method is purposed to reduce the cost of Web document automatic processing and improve processing efficiency by providing information about the structure of Web pages using templates through applying the method to information retrieval such as data extraction.

International Journal of Contents

A Study on Effective Internet Data Extraction through Layout Detection

Abstract

Keywords