Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I think the point is that it's already been digitized. All of it. The stuff in your attic isn't a valuable historic artifact; it's deep in the trough of no value.


A fraction has been digitized, and most of it still quite poorly scanned.

Try reading some, try actually using any of the schematics and other technucal drawings, a lot looks ok at a glance, then turns out to be merely better tgan nothing when you actually try to use it the same way you would have used the original.

Every few years the definition of practical scanner quality and file size increases by 2x or more. A fax quality scan is infinitely better than nothing, so it's great that such scans were made as soon as that became possible, but a few years later scanners became 10x better and it became practical to work with 10x larger file sizes. So all of those documents really need to be scanned again. And the same process needs to hapoen yey again at least one more time. Even very good scans from only a few years ago are still only good compared to old scans. They are still practically garbage compared to the original print.

And much of what is currently digitized is precarious, one lawsuit or lapse in funding or ceo decision away from the service being shut down.

I've already lost tons of stuff I actually paid for to services that no longer exist, let alone free services.

Right now a lot of eggs are in archive.org and they are a charity that paints a huge "kill me" target on their own back every day. It makes no sense to operate on the assumption they will never become the library of alexandria.

Digitizing existing printed material is not remotely a done job. Not even barely scratched the surface if you step back and look at how we are still in the first few seconds of eternity.


You'd be surprised how little of the niche pre-internet stuff is digitized. There are some exceptions (national-circulation newspapers, some English-language books), but it's a minority of what's been published.


People have absolutely no idea how niche things were - every city of moderate size (maybe a million?) had an independent computer magazine of some sort, some literally mimeographed pages stapled together, others flashy magazine quality productions.

And they had columnists and journalists and everything. Heady time.


And I'm not sure I can go online and find a complete set of Byte Magazine much less the Boston Computer Society newsletter or something even more niche.

Probably I could find a lot of stuff if I went to enough libraries but it certainly isn't all available and nicely indexed on the open web.


Also a question of how long stuff will really stay around on the internet.


Depends how many FTP sites it's mirrored on.


A good third of HP Application Notes is still not digitized. Same with many other publications that aren't Dr. Dobbs' or BYTE.


How many backups are there of it? Recall that kernel.org was once completely lost and much of the source code apparently had no mirrors.


I may have the only omnibus of PET Benelux Exchange magazines out there, which seem to be entirely unavailable online. There are exceptions.


Scan and share!


Yeah, I had grand plans. I even did one issue, and OCR'd it as well, but then COVID happened and work got busy and now you can fast-forward four years when I've moved countries again, hopefully for the last time. I know where they are, which is good. Need to disassemble the book-edge scanner and clean the underside of the glass off because it's all fogged. I also need to stand up a VM that can talk to the crappy OpticBook drivers. And I'm behind on my start-of-season farm work, which keeps expanding fractally the deeper I go. That dependency stack is what's in my way at the moment.

Eventually, I swear!


Exactly that. The major magazines have been done already.

If you have something obscure, there's a good chance it hasn't.


In which languages? And who has access to the digital copies? And how complete? This isn't as clear cut or solved as you make it out to be.


This is why I data hoard

I have friends who are Serbian living elsewhere and are about to start a family

They have no Serbian children’s media afaik (and it’s not exactly easy to find)

I found ~200GB of Serbian dubbed children’s media someone painstakingly digitised on /t/

I’m keeping it for when they eventually have those kids

Unless shown otherwise, assume nothing is archived correctly, completely or at all


What will happen to your hoard after you are gone?


Hopefully copyright on most of it lapses so I can freely upload to Archive.org or similar


Have you thought about putting it on a torrent or sharing it with r/datahoarder ?


It’s still up on /t/ afaik




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: