yourgeek.it_

~/ never-deploy-on-friday

He Added a Code Snippet 20 Minutes Before a Demo

#incident#troubleshooting#infrastructure#human-factor

My phone rang on a Friday. “Help me, Obi-Wan Kenobi. You’re my only hope.”

It was the webmaster of a WordPress site hosted on a server I manage. The site was down, and he had a demo in twenty minutes.

Minutes, not hours

I logged in over SSH and read the php-fpm log. It said the same thing over and over: it could not allocate more memory.

Then I called him and asked the only question that matters: “Did you change anything lately?” He said no, nothing, only a plugin. It lets you add snippets of code to WordPress without touching the theme, and he had used it to add one that restricted access to some sections of the site.

I renamed the plugin’s directory. WordPress noticed the plugin had vanished and deactivated it by itself, and the snippet went with it. The site came back, and the demo went very well.

From phone to live again, it took a few minutes.

I never read the snippet. My deduction: it ran an infinite loop, so every request kept asking for memory until it hit the ceiling. I haven’t proved it, and I didn’t dig further, because the thing that mattered was already fixed.

Why nobody else noticed

The server is shared, so every site runs in its own php-fpm pool with a memory limit. Each request from the broken site died when it hit the ceiling. The other sites on the same machine never noticed anything.

Without that limit, one site’s loop could have eaten the whole box.

Same plugin, 18:05

Now move the same change to Friday at 18:05.

This time the webmaster installs the plugin and adds the snippet at 18:05, and I am going to have an afterwork drink, the Friday kind, the one that decompresses the week.

The bug is identical. The log says the same thing. But nobody is looking at it. The webmaster sees the site go down and does what anyone would do: he calls, and calls again, increasingly desperate, because on Friday evening nobody can do anything. The site stays dead until Monday morning. Panic makes it worse: I wrote about what it pushes you to do in Don’t panic: the order matters more than the speed.

The failure was never the snippet. It was the moment: a change nobody could undo, with nobody reachable.

What I’d tell you

The webmaster in a small company is often also sales and support. A demo in twenty minutes is a good reason to change something, and a bad reason to change it on production. Those are the people I’d like to reach, and also the sysadmins who prepare their environments.

  • Staging is mandatory. If the webmaster doesn’t want one on the server, he needs a sandbox on his own machine. A new plugin, and any snippet of code, goes there first, always.
  • A change that looks trivial on Friday afternoon is the one to distrust.
  • If you can’t undo it with a single command, or nobody is reachable afterwards, it waits until Monday.
  • Put a memory limit on every site. The damage stays inside the site that caused it.
  • Know who answers the phone, make sure they can reach the server, and make sure they know what to do once they’re there. Some people get in and then ask “now what?”

← all posts