Probe, Count, and Classify: Categorizing Hidden-Web Databases

Luis Gravano
Panagiotis Ipeirotis
Mehran Sahami

Venue: Proceedings of the 2001 ACM International Conference on Management of Data (SIGMOD), 2001
May 2001
Status: Refereed
Type: Conference
Acceptance rates: 15% accepted

The contents of many valuable web-accessible databases are only accessible through search interfaces and are hence invisible to traditional web “crawlers”. Recent studies have estimated the size of this “hidden web” to be 500 billion pages, while the size of the “crawlable” web is only an estimated two billion pages. Recently, commercial web sites have started to manually organize web-accessible databases into Yahoo!-like hierarchical classification schemes. In this paper, we introduce a method for automating this classification process by using a small number of query probes. To classify a database, our algorithm does not retrieve or inspect any documents or pages from the database, but rather just exploits the number of matches that each query probe generates at the database in question. We have conducted an extensive experimental evaluation of our technique over collections of real documents, including over one hundred web-accessible databases. Our experiments show that our system has low overhead and achieves high classification accuracy across a variety of databases.

QProber: Classifying and Searching “Hidden-Web” Text Databases

Panos Ipeirotis

Probe, Count, and Classify: Categorizing Hidden-Web Databases

Panos Ipeirotis

Probe, Count, and Classify: Categorizing Hidden-Web Databases

Related Files:

Related Projects:

Panos Ipeirotis