Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Google's cache has approximate 480 story pages still in its cache. The Wayback Machine won't show any results for the site because of its current robots.txt. It'd be great if someone could scrape these pages and provide them as an archive. I am not sure how to do that without running afoul of Google's terms-of-service.

List of URLs: https://dpaste.de/YbEq/raw



I would love a copy of that! Since I have diaries here: http://k5.semantic-db.org/diary-slurp/161942--archive-diarie...

I regret not scraping the stories too.


You can use http://archive.is/



(Update: I put what I managed to salvage up at http://atdt.freeshell.org/k5/).


I was hoping "What Art Is" would be salvageable, but sadly not.

  http://www.kuro5hin.org/story/2005/6/19/143553/870
I don't even recall who wrote it now, but I've had it bookmarked for over 10 years.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: