I'm currently an undergrad in Computer Engineering. As an engineer we are taught to look ahead when designing software and to do it with the "good of society" in mind.
I find it VERY hard to believe that Facebook did not plan ahead and implement an efficient method of photo deletion. If they truly did not, I seriously call upon their skills and insight in creating software for the public.
Come on Facebook ... it's deleting a damn image! This should be top priority in the modern age of personal privacy.
Honestly, I'm not so sure. Looking at the APIs exposed by Facebook you'd be hard pressed to find anything that looks like its engineered. It seems to be a hodge podge of features thrown together over years with the stuff that works best/well enough sticking around long enough to be seen by the public. I think its a safe bet that many of the things not exposed are quite messy/nonsensical.
They could very well have never thought of deleting a photo; it doesn't seem that there is much interest in information removal there, and I could easily see if being a very ugly manual process that no one wants to spend time doing and as such it never gets done.
Note that the delete operation merely marks the data deleted - they don't actually purge or overwrite it. Perhaps there's some kind of bug in this deleted-marking system? Kind of weird they haven't been able to fix it yet though.
As an engineer we are taught to look ahead when designing software and to do it with the "good of society" in mind.
Keep in mind, that when the photo features of Facebook were added, they probably didn't do it while thinking about half a billion users or whatever gazillion terabytes of photos they have to operate with. Heck, they may have even started out with simply storing photos on disk in their servers where deletion was easy as opposed to on a CDN.
I've designed a couple of small scale websites before and I've always thought about how they would react if the user base would expand significantly. It's always something in the back of my mind when programming.
Even with a small user base; if a user flags a photo for deletion -- it should get deleted. Period.
1. Delete it from the public facing servers that we control.
2. Put a request in the queue so the backups, and thumbnails will be deleted.
3. Pray that the CDN will purge the picture eventually.
#3 is a big deal... we don't control the CDNs cache options, and can't push a 'delete' upstream so the picture might be gone 2 seconds after the deletion is completed on our end, or it might take 2 weeks.
I'd thought Facebook would have been able to engineer or negotiate better options than we do, but apparently not.
There's large scale, there's really large scale and then there's Facebook scale. They operate at a scale that almost no one does, and in fact, if you're thinking of how to operate at that volume when you start your new project, then I'd say you're doing it wrong (I know you didn't actually say that - just pointing out that no one should be thinking of how to operate at that scale when designing their basic architecture).
I don't really know how their photo system works - maybe it is a trivial thing to purge files from CDNs on demand - but I'm willing to give the benefit of the doubt and say never attribute to malice, what could be explained by earth-population-level scaling problems.
Also unrelated: Facebook has shown many times that they can't really be trusted on privacy issues - I never upload anything there under the illusion that I can freely delete it and it's really gone. But I don't personally believe this thing is a privacy issue they're intentionally balking on
Thanks to those of you who give Facebook the benefit of doubt. I don't work on this particular system, so I'm not qualified to discuss any specifics, nor do I speak for the company. I can say this though: People @ Facebook do care about these types of problems and work to solve them. I'm not some new grad; I've worked with large systems for a while. The scale is hard for me to comprehend. The systems I do work with illustrate very well that nothing is as simple as you think it might be and most of the obvious solutions get tossed out the window because they won't scale.
Maybe I mis-represented my argument. I didn't mean that I code small scale projects with large scale back ends just in case a user base boom happens. I meant that I try to code as using a "module" approach as much as possible. When (and if) the opportunity presents itself, the small scale project will be able to integrate a large scale solution much more easily.
You're making the same mistake that many awful managers make; namely "If the problem seems like it should be easy, it must be easy".
You aren't familiar with Facebook's infrastructure and comparing working on small websites to working on something that scales to roughly 1/12 of the entire population of the planet is laughably naive.
Saying you use a "modul[ar]" approach is nothing but handwaving. You think that facebook doesn't use modules? Or doesn't have a service oriented architecture to some degree? They friggin' developed their own RPC transport!
I'm currently an undergrad in Computer Engineering. As an engineer we are taught to look ahead when designing software and to do it with the "good of society" in mind.
Don't worry a few years in Industry and you'll learn what things are really like.
i.e. There is no end of shite software that wasn't developed. Yes you can make lots of money with hodgepodge software. Welcome to Industry :)
This kind of cache (mapping a url to a constant image) seems to have little to do with that kind of cache (mapping a memory address to likely changing memory contents.)
Can you clarify why you think the problems are similar?
If you start from the premise that there are only two hard CS problems then all hard CS problems must be a special case of one or both of those two. So cache in an on-chip memory sense is conflated with disk storage is conflated with distributed disk storage. This isn't necessarily bad, Smarty for example caches rendered files to disk storage and calls it a cache, but it does show off the problem of naming things. (That and every variable named "data".) Anyway, I imagine what the GP was getting at is that if your data distribution isn't particularly deterministic enough (e.g. swapping hard drives in and out when they fail) you have to deal with validation that a particular data changing command (store, delete, whatever) actually propagated to a sufficient portion (in some cases all) of the servers and that the introduction of new or changing or rogue servers doesn't affect that. The more apt term for this is consistency. Related is the CAP theorem which says for any distributed system, you can only pick any 2 of consistency, availability, or partition tolerance, though it's more interesting to talk about atomic operations, where transactions are/where they might be desired, and whether things get faster or slower with more data.
They're similar in that a deletion consists of a change to the image. This means that the image is now changing, so it's the same basic problem again.
The two scenarios have entirely different constraints, this is undeniable. It's still an instance of the same basic problem: verifyably deleting the image means invalidating all caches of that image.
I find it VERY hard to believe that Facebook did not plan ahead and implement an efficient method of photo deletion. If they truly did not, I seriously call upon their skills and insight in creating software for the public.
Come on Facebook ... it's deleting a damn image! This should be top priority in the modern age of personal privacy.