Impossible. The majority of websites firewall automated crawler traffic (because of the rise of the bots), only making exceptions for the largest search engines. There is no possibility of starting a new crawler.
It would be interesting to see a decentralised, residential collective that builds and publishes an index. There are surely enough interested people on HN alone that would be willing to run software at home to scrape a small slice of the internet.
The majority of websites try to do that but they do not catch as much traffic as they think they do. A starting point for a scraper is to run it on your home connection in an undetectable web driver framework such as zendriver.
This was a knee jerk response to the first paragraph. They weren't talking about a general crawler, but a subset of Wikipedia, stack overflow, programming docs and github. You can download archives of all of those except github, and github could be queried using the api or GH cli