HIT campus search engine system which clusters your search results into topics. The user will quickly find what you are looking for. This system includes four parts: 1. Web crawling 2. HTML parsing 3. Indexing 4. Searching. We use the open source software "Hetrix" as our cralwer, use the Lucene to build the index. In order to quickly find what you are looking for, we use carrot2 to help us cluster the search results into topics. We also write a script to fetch the websites in the campus everyday and update the index automatically.
License
W3C LicenseFollow HITSearchEngine
You Might Also Like
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of HITSearchEngine!