Focused Crawling for both Topical Relevance and Quality of Medical Information

Tang, Tim; Hawking, David; Craswell, Nick; Griffiths, Kathleen

Focused Crawling for both Topical Relevance and Quality of Medical Information

Date

2005

Authors

Tang, Tim

Hawking, David

Craswell, Nick

Griffiths, Kathleen

Publisher

Association for Computing Machinery Inc (ACM)

Abstract

Subject-specific search facilities on health sites are usually built using manual inclusion and exclusion rules. These can be expensive to maintain and often provide incomplete coverage of Web resources. On the other hand, health information obtained through whole-of-Web search may not be scientifically based and can be potentially harmful. To address problems of cost, coverage and quality, we built a focused crawler for the mental health topic of depression, which was able to selectively fetch higher quality relevant information. We found that the relevance of unfetched pages can be predicted based on link anchor context, but the quality cannot. We therefore estimated quality of the entire linking page, using a learned IR-style query of weighted single words and word pairs, and used this to predict the quality of its links. The overall crawler priority was determined by the product of link relevance and source quality. We evaluated our crawler against baseline crawls using both relevance judgments and objective site quality scores obtained using an evidence-based rating scale. Both a relevance focused crawler and the quality focused crawler retrieved twice as many relevant pages as a breadth-first control. The quality focused crawler was quite effective in reducing the amount of low quality material fetched while crawling more high quality content, relative to the relevance focused crawler. Analysis suggests that quality of content might be improved by post-filtering a very big breadth-first crawl, at the cost of substantially increased network traffic.

Keywords

Keywords: Costs; Data reduction; Information retrieval; Search engines; Telecommunication traffic; World Wide Web; Domain-specific search; Focused crawling; Quality health search; Health care Domain-specific search; Focused crawling; Quality health search

URI

http://hdl.handle.net/1885/83541

Collections

ANU Research Publications

Source

Proceedings of ACM Conference on Information and Knowledge Management (CIKM 2005)

Type

Conference paper

Restricted until

2037-12-31

Downloads

File

Description

01_Tang_Focused_Crawling_for_both_2005.pdf (160.75 KB)

Full item page

Cultural advice

Focused Crawling for both Topical Relevance and Quality of Medical Information

Date

Authors

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Description

Keywords

Citation

URI

Collections

Source

Type

Book Title

Entity type

Access Statement

License Rights

DOI

Restricted until

Downloads