Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Just an FYI if you're not aware: the entire corpus of reddit comments, users, submissions and subreddits is publicly available as a dataset from various sources. There is a very straightforward (and as of now, loosely sanctioned) method of continually crawling all reddit activity that doesn't rely on the API. Judging by your complaints you might find this useful.

But I agree that reddit is likely restrict access even more in the future.



Would you have more info on this? I’d like to make sure the entirety of Reddit makes it into the Internet Archive before they go full on Digg.


http://files.pushshift.io/reddit/

I have a little open source project for rudimentary indexing and search of the dataset as well as adding sentiment data: https://github.com/dewarim/reddit-data-tools




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: