PrisCrawler: A Relevance Based Crawler for Automated Data Classification from Bulletin Board

Computer Science – Information Retrieval

Scientific paper

Rate now

  [ 0.00 ] – not rated yet Voters 0   Comments 0

Details

published in GCIS of IEEE WRI '09

Scientific paper

Nowadays people realize that it is difficult to find information simply and quickly on the bulletin boards. In order to solve this problem, people propose the concept of bulletin board search engine. This paper describes the priscrawler system, a subsystem of the bulletin board search engine, which can automatically crawl and add the relevance to the classified attachments of the bulletin board. Priscrawler utilizes Attachrank algorithm to generate the relevance between webpages and attachments and then turns bulletin board into clear classified and associated databases, making the search for attachments greatly simplified. Moreover, it can effectively reduce the complexity of pretreatment subsystem and retrieval subsystem and improve the search precision. We provide experimental results to demonstrate the efficacy of the priscrawler.

No associations

LandOfFree

Say what you really think

Search LandOfFree.com for scientists and scientific papers. Rate them and share your experience with other people.

Rating

PrisCrawler: A Relevance Based Crawler for Automated Data Classification from Bulletin Board does not yet have a rating. At this time, there are no reviews or comments for this scientific paper.

If you have personal experience with PrisCrawler: A Relevance Based Crawler for Automated Data Classification from Bulletin Board, we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and PrisCrawler: A Relevance Based Crawler for Automated Data Classification from Bulletin Board will most certainly appreciate the feedback.

Rate now

     

Profile ID: LFWR-SCP-O-564419

  Search
All data on this website is collected from public sources. Our data reflects the most accurate information available at the time of publication.