What About Existing Open Source Crawlers?
A few years in the past, Creative Commons tasked me with building a web crawler able to downloading 500 million photos. Crawling something beyond a number of thousand URLs calls for a quick distributed system. Moreover, it is not sufficient to be fast; moral, authorized, and practical concerns demand that a crawler be polite: a […]
What About Existing Open Source Crawlers? Read More ยป
