A PHP process in /tmp
The alert went red: an uptime monitor was reporting read timeouts on an e-commerce site. The shop belonged to the company I worked for. A third party ran it, and we only provided the hosting.
I went looking for why PHP had stopped answering. Among the processes I found one that had nothing to do with the shop’s source tree: a PHP process running from /tmp, handling incoming traffic on its own.
My first thought was to kill it. It was choking the server, it had no business being there, and killing it would have taken a second. It would also have erased the only evidence of how the intrusion worked. I didn’t do it. I stopped, and sat down with my AI assistants to work out a plan before touching anything.
Isolate, preserve, clean, verify
The plan had four steps, in this order: cut the machine off from the network, preserve the evidence, clean, verify.
From the alert to the isolation took minutes. Then I took a dump of the process and put it in a tar archive together with the logs and the patched files, so the evidence would still be there for later. Cleaning took about half an hour, then a quarter of an hour of checks, and the site went live again.
The timestamps put the attacker inside for a few hours before the alert. The cause was a known vulnerability in an e-commerce platform that hadn’t been patched for months. The company running the shop had flagged the patch; it was stuck behind commercial agreements. A very ordinary way to end up here.
What the intruder left behind was a process redirecting visitors to a third-party site. The risk was not lost data but customers typing their credentials into a fake shop. The checks we ran afterwards found no sign of credentials being entered during the window of the breach.
No data left the server, for two reasons, both by design. The database ran on a different host. And the CI/CD pipeline deploys configuration files with placeholders only: the real values are injected on the server at runtime. An attacker reading the config gets placeholders, not access.
Had I killed the process, I’d have had a clean-looking server and no answer to the two questions that matter: how did they get in, and what did they touch?
The one I caused
Some time before, I had wiped a mail server’s disk myself. I was merging two VMs by hand, on the same cluster that ran production, one moment of distraction. The whole story is in the Proxmox post. Here is the part that matters for this one.
The moment I understood what I’d done, I panicked. My first thought was to restore the disk, right now. I was about to do it when I did the arithmetic: the restore would take long, and for all that time nobody would have email.
So I brought up a clone of the server instead. Configuration first, then webmail, with IMAP and SMTP closed to clients. People used webmail while the mailboxes came back from backup. I followed the logs with tail and sent test messages from an external mail server, because from the inside everything always looks fine. About an hour of work, a few minutes of visible downtime, no mail lost.
What panic does
It was the same reaction both times, with the urge pointing in opposite directions.
On the shop, the urge was to destroy: kill the process, wipe the machine, make the intruder disappear. On the mail server it was to repair fast: restore the disk, make the mistake disappear before anyone noticed. Anger at an attacker pushes you to stamp it out; guilt pushes you to undo it. In both cases the action feels right, and in both cases it was the right action at the wrong moment.
I didn’t get rid of the panic either time. It was there from the first second. What worked was letting it run in the background while I asked two questions before every action: what does this destroy, and what does it cost in time? Killing the process destroys the evidence. Restoring the disk means hours without email.
The Hitchhiker’s Guide to the Galaxy carries those two words on its cover, in large friendly letters, for readers who are about to have a very bad day. I’ve always had them in mind during an incident: an instruction for what to do while you’re not calm, not advice to be calm.
Know where the valve is
In the post about the water filter the step that saved the floor was knowing where the main valve was before I needed it. Incidents work the same way. Panic doesn’t leave room to invent things, so what you have ready decides how the first minutes go.
Looking back at the two cases, this is what was already in place, and what I’d make sure of on any system I run:
- Something that tells you, from outside. The shop incident started with an uptime monitor, not a customer.
- A way to cut a machine off from the network without losing your own access to it.
- A place to put evidence before you clean anything: a dump, the logs, the changed files, in an archive.
- A test that doesn’t run on the thing that is broken. A message sent from an external mail server proves more than a green light inside.
- A plan, even a rough one, written down before the first command. Mine took a few minutes with my assistants, and it kept my hands off the kill switch.
The shop had one more thing that no runbook covers: the patch had been flagged and was stuck behind commercial agreements. Today I’d want patching written into the hosting agreement with whoever runs the application.
Takeaways
- Panic is normal. Let it run in the background and work the sequence.
- Order: trace, contain, preserve, clean, verify, repair.
- Before every action ask: what does this destroy, and what does it cost in time?
- The first instinct is the right action at the wrong moment.
- Don’t kill the process before you’ve dumped it.
- Test from outside, never only from the machine that broke.
- Whether the incident is yours or someone else’s, the order is the same.
- Find out where the valve is before the flood.