Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I've had a deployment cycle like this since I started using Google App Engine... 6 years ago.


The interesting thing about the article isn't that they're able to release continuously; there's nothing technically hard about deploying quickly. The interesting thing is that they're able to make the continuous release system work with thousands of engineers actively working in the same codebase without destroying quality. The three-tiered release system, monitoring alerts, feature flags, and good testing infrastructure seem to be what makes all of that possible.


Exactly - releasing is the easiest part of the process.

A lot of orgs don't have continuous deployment because of reasons such as:

- they don't have a good enough automated testing suite (or at least don't trust it fully), and thus rely on "sign offs" to have people commit to saying it's quality

- they don't measure in production properly (no real error alerts, no way to measure release success), and often deal with things in a "go or no-go" type way

- they don't canary test. To me this one is critical - the only way to get real production use is to have real production users actually using the site/platform/app, just a sample of them, to see what could go wrong, especially with new features

A lot of managers I've worked with are shocked whenever I pull out the "continuous deployment is easy. doing it well is hard" line.


Yup, we do all that, except the 10k developers scale. Ha.

I had the advantage of starting with fresh codebases and a small team. Obviously, adding this to an existing organization is much more difficult.

When you setup your system correctly from the start, it also becomes a great hiring tool. Once you show developers the environment you work in, mouths drop and they almost beg to work for you.

Doing it is hard, but not impossible. At this point there really is no excuse to not start things off like this. It is about 1-2 weeks of effort to setup a new project with all the right tooling on top of GCP thanks to the features they give you as part of their platform.


I find a big thing too is building up that mindset of always roll forward. Bug in production means next production build will fix said bug. No hotfixing to say the release is done etc.

Completely agree that from a tech side there is no excuse, however a lot of QA culture has persisted through orgs and they want to keep that feeling of control (even though automation does it way better than them)


A lot of organization don't have continuous deployment because they can't risk breaking everything for any developers who is playing around. When they want to release, they review everything, test and go through QA.

Facebook is not important. It has no impact when it's broken.


I helped build a business that did about $80m in gross revenue in the first year. We launched the initial version in 3 months (which we predicted to within a week).

Started with 2 engineers (myself and another guy) and grew it to about 15. Zero QA, Zero DevOps.

We had CI/CD and a full test suite. We deployed from master as many times a day as we needed / wanted.

It can work if you open your mind to it and you hire the right people who know what they are doing.


And I was at business with $800M and 30 employees.

Just because it releases quickly and has no QA doesn't mean it's a good thing.

The only metrics that matters is calls from your users. Facebook doesn't even have a number to call when it's broken.


Keep in mind that the quarterly earnings report was 9.32 billion USD. That is approximately 100 million a day. I'd say downtime at that scale is important.

Also: People tend to call the police: http://time.com/3071049/facebook-down-police/


It's extremely low revenues per person and per page view.


Except to their own companies bottom line, generating millions in revenue?

Releasing less often is a way to guarantee that bigger bugs will get through at some point, requiring hotfixes etc. The more you release, the higher quality releases you have, and the smaller production incidents.

The point isn't to remove QA, it's to trust that the automation in place is high quality and will catch the majority of issues before they are issues (i.e. if something makes it through the automation, then it should be caught in the internal release, or at least the 3% canary group), and then on the back of issues, make the automation more robust.

The more people who are introduced into a process, the more likely it is to fail at some point - the fact is that Facebook has a pretty low rate of huge production issues compared to most software companies, they must be doing something right.


Whereas I'm sure whatever you worked on had 1.32 billion active users everyday.

Facebook is "7th most valuable company in the world" important.


1.32 billion FREE users.


1.32 billion units of inventory

FTFY




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: