As this article points out, it's tremendously unclear who is using residential proxies.
The big AI models claim they're not using them. I'm not inclined to "just believe them", but no incriminating evidence has leaked, and—as pointed out in the article—many of the bots that are running on these residential proxy botnets are coded in incredibly stupid and inefficient ways.
How confident are people who research this stuff that the RP botnets are actually being used for AI training?
That's the only theory we have that doesn't sound like a conspiracy theory. The only other credible ideas are that someone's doing some kind of dataset arbitrage by scraping the fuck out of everything and selling companies data that is technically new (by means of the scrape date being newer).
Right, and as I understand it the timing also lines up: these ill-behaved scraper-bot-nets exploded along with GenAI in the last 3-4 years.
Still, it seems to me like something doesn't add up. Running these botnets is perhaps cheap but it isn't free, and dumping all their data into LLM training is truly expensive: would so many of these bots be so blitheringly inefficient in their scraping patterns if all of the results were getting fed into LLM training?
Have there been no leaks or whistleblowers from the "semi-legit" RP brokers?
The big AI models claim they're not using them. I'm not inclined to "just believe them", but no incriminating evidence has leaked, and—as pointed out in the article—many of the bots that are running on these residential proxy botnets are coded in incredibly stupid and inefficient ways.
How confident are people who research this stuff that the RP botnets are actually being used for AI training?