Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

My immediate thinking was this as well: when they stop touching it, it's fine, which means that whatever bugs are checked in seem to be caught quickly. What's more, the periods of high reliability are during periods when you might expect there to be heavy load - assuming people don't check Facebook at work (har har).

What counts as an incident?

> Figure 1 includes data from an analysis of the timing of events severe enough to be considered an SLA (service-level agreement) violation. Each violation indicates an instance where our internal reliability goals were not met and caused an alert to be generated. Because our goals are strict most of these incidents are minor and not noticeable to users of the site.

So these are minor issues. The parent article is paraphrasing at best, and jumping to conclusions at worst. From Facebook:

> We believe this is not a result of carelessness on the part of people making changes but rather evidence that our infrastructure is largely self-healing in the face of non-human causes of errors such as machine failure.

They then list a number of sane strategies to mitigate this.

http://queue.acm.org/detail.cfm?id=2839461



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: