Study on Distance Measures for Clustering of Web Documents based on DOM-Tree based Representation of Web Document Structure

Main Article Content

Manoj Kumar Sarma, Anjana Kakoti Mahanta

Abstract

Among the three broad areas of Web mining, Web Structure Mining is the method of discovering structure information from either the web hyperlink structure or the web page structure. In order to apply data mining techniques on web pages, a good and efficient representation of web pages is required that could depict the actual hierarchical structure of web pages. The work presented here aims to find out an appropriate distance measure (also called as similarity measure) for strings that can be used for clustering of web documents and also for other data mining applications.

Article Details

How to Cite
, M. K. S. A. K. M. (2017). Study on Distance Measures for Clustering of Web Documents based on DOM-Tree based Representation of Web Document Structure. International Journal on Recent and Innovation Trends in Computing and Communication, 5(6), 1440 –. https://doi.org/10.17762/ijritcc.v5i6.972
Section
Articles