Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

What about tor.


TOR is an easy one to detect. Case in point: Try to browse the web with TOR and see how many CloudFlare captchas you have to solve in the first 10 minutes ;)

Also TOR has public lists of exit nodes (https://torstatus.blutmagie.de/), and unless you exit via a non-listed one you're trivially identifiable.

Lastly, TOR is rather slow, which prohibits using for large-scale scraping tasks. Plus you get switched around different exits and countries, which might break your scraping logic.


I believe you could use the ExitNodes option to only use a specific exit.

There's no such thing as an unlisted exit node, only unlisted entry nodes (bridges). I agree that Tor isn't a good choice for scraping.


Tor would introduce pretty intolerable latency into a professional data crawling setup.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: