TOR is an easy one to detect. Case in point: Try to browse the web with TOR and see how many CloudFlare captchas you have to solve in the first 10 minutes ;)
Also TOR has public lists of exit nodes (https://torstatus.blutmagie.de/), and unless you exit via a non-listed one you're trivially identifiable.
Lastly, TOR is rather slow, which prohibits using for large-scale scraping tasks. Plus you get switched around different exits and countries, which might break your scraping logic.