+− THE DAILY DIFFdev & AI news
NEEDS REVIEW

Facebook deleted itself from the internet. Six hours.

Facebook removes its own address from the internet, and when its engineers arrive to put it back, the badge readers are down too, because they run on Facebook.

Facebook removes its own address from the internet, and when its engineers arrive to put it back, the badge readers are down too, because they run on Facebook. October 4, 2021, 15:40 UTC: a routine backbone-maintenance command takes down every link between Facebook's data centres; the audit tool built to block it has a bug; Facebook's name servers, unable to reach the data centres, withdraw their own BGP routes by design — and facebook.com, Instagram, WhatsApp and Messenger vanish from DNS for nearly six hours. Internal tools, out-of-band access and the office badges ride the same network, so the fix needs people at the racks.

Read the written edition (English) ↗

What this video covers

  • Facebook removes its own address
  • The receipts: Meta Engineering, Cloudflare, Hacker News
  • This is The Daily Diff, postmortem
  • Timeline: 15:39 command → 15:40 withdrawn → 21:20 back
  • Mechanism: the health check that deletes you, the 30× DNS storm, the doors

Transcript

Facebook removes its own address

0:00 Facebook removes its own address from the internet, and when its engineers arrive to put it back, the badge readers are down too, because they run on Facebook. Meta's own postmortem, next day: a routine maintenance command takes down every backbone link, and the tool built to block commands like that has a bug.

The receipts: Meta Engineering, Cloudflare, Hacker News

0:16 Cloudflare, from outside: at 15:40 UTC the routes to Facebook's name servers vanish, and dig facebook.com returns SERVFAIL everywhere on the planet. How it happens, why a safety feature makes it worse, and who gets the blame. This is The Daily Diff, postmortem. Early afternoon, UTC.

This is The Daily Diff, postmortem

0:34 A command meant to check free backbone capacity instead takes down every link between Facebook's data centres, and the audit tool built to catch exactly this

Timeline: 15:39 command → 15:40 withdrawn → 21:20 back

0:41 waves it through. 15:40. Cloudflare's BGP feed fills with withdrawals; Facebook's DNS prefixes leave the routing table. Within a minute Cloudflare's engineers are in a room wondering whether they broke 1.1.1.1. 15:45, Hacker News, twenty-six hundred points. 16:07, Facebook's spokesman, on Twitter of all places: some people are having trouble accessing our apps. Some people is three and a half billion.

1:06 18:51, the New York Times: employees locked out of their buildings, badges dead. 19:52, the CTO apologises. 21:00, five hours in, routes return; 21:20, facebook.com resolves. Why does a backbone fault delete Facebook from the internet?

Mechanism: the health check that deletes you, the 30× DNS storm, the doors

1:20 Its name servers have a rule: can't reach the data centres, declare yourself unhealthy, stop advertising your address. Sensible when one site goes bad. When the whole backbone goes, every name server does it at once. The servers are up. Nobody can find them. A safety feature turns a network fault into an existence fault. Then the internet piles on: apps retry, users reload,

1:39 and Cloudflare's resolver sees thirty times its normal load, for a site that isn't there. And the doors. Tools, out-of-band access, badges: same network. Engineers drive to the data centres, can't get in, then meet racks hardened against anyone with physical access. Built to slow an attacker, it slows the owner too. The backbone returns; they ramp traffic slowly, on purpose.

1:59 Each data centre has dropped tens of megawatts, and reversing that in one step is a second outage.

git blame — the split

2:05 git blame. Facebook, fifty-five percent: the command was audited, the auditor had the bug, one command reached every router. The DNS design, twenty-five: a health check with no sense of proportion. One network, fifteen: tools, out-of-band, badges. The racks, five, for doing their job. Blast radius: three and a half billion users, nearly six hours. The stock closes down five percent, six billion off Zuckerberg,

Blast radius

2:27 on paper. Telegram claims seventy million sign-ups that day. Verdict, postmortem: needs review. Same-evening statement, named-author postmortem within a day, audit-tool bug admitted in plain English.

Verdict + the Monday line

2:38 The fix list: strengthen testing and drills. Nothing about taking the doors off the production network. Monday: put your out-of-band access, and your door, on a network you don't operate. Send me the incident you're still not allowed to talk about, in the comments, or at the daily diff dot dev. And that's the diff for today. I'm Niko from Axrisi.

2:58 Merge responsibly.

Sources

  1. Meta Engineering, "More details about the October 4 outage" (Santosh Janardhan, Oct 5, 2021)engineering.fb.com
  2. Meta Engineering, "Update about the October 4th outage" (Oct 4, 2021)engineering.fb.com
  3. Cloudflare, "Understanding how Facebook disappeared from the Internet" (Oct 4, 2021)blog.cloudflare.com
  4. Mark Zuckerberg, note to employees (Oct 5, 2021)www.facebook.com
  5. Andy Stone (Facebook), 16:07 UTCtwitter.com
  6. Sheera Frenkel (NYT), 18:51 UTC, the badgestwitter.com
  7. Mike Schroepfer (Facebook CTO), 19:52 UTCtwitter.com
  8. Hacker News (2,589 points)news.ycombinator.com
  9. The New York Times, "Gone in Minutes, Out for Hours"www.nytimes.com
  10. Forbes, "Zuckerberg Loses $5.9B in a Day" (estimate)www.forbes.com
  11. Reuters, Telegram's 70 million sign-ups (Durov's figure)www.reuters.com
  12. Krebs on Securitykrebsonsecurity.com

Related videos