Internet Archive's open-source, web-scale, web crawler project
Free batch downloader for image, wallpaper, video, audio, document,
Download websites as e-book: pdf, txt, epub.
Open source web crawler for Java
Hadoop framework for scalable processing of large web corpora
Capable to "Crawl" a site and return a report of all links from it
Combined search engine for publication databases.
Crawls reddit website to pull statistical info.
A Lua-based crawling scripting language and leveraging selenium