yourgeek.it_

~/ full-stack-on-1gb-ram

A full stack on 1 GB of RAM

#infrastructure#resilience#troubleshooting

One small box, on purpose

The server that runs my identity provider, my mail server, my database and my API has 1 GB of RAM. Right now the whole stack uses a little over half of it.

It is a single machine, public on the internet. Traefik is the only ingress, on port 443, which means every service behind it has to find the real client IP in a header, not on the socket. Behind it sit Authentik for single sign-on, Postgres, an API written in Rust, and a mail server with spam filtering, plus a small container that renews its certificate. Seven containers in total.

This is not a scale story. Most of what runs there serves my own projects, and I am the only person who logs in. It is a story about a habit: deciding how much a service gets before it starts, then checking what it really takes.

I picked the small machine on purpose. A size I can’t hide behind.

It is not the first time I run something important on very little redundancy: for seven years I ran a company’s internal infrastructure on a three-node cluster without high availability, which I wrote about in Seven Years of Proxmox in Production, Without HA.

A limit makes you optimize

The reasoning has nothing to do with saving money. On a big machine there is always the thought “there’s plenty of hardware, use it”. Nobody measures, because nothing forces them to. A small machine forces the question: how much does this service need?

So nearly every service gets a memory ceiling before it runs, in the Compose file:

services:
  api:
    deploy:
      resources:
        limits:
          memory: 128M

The ceiling is the budget, docker stats is the measurement. When the two disagree, one of them is wrong, and I would rather find out while nothing is on fire.

Raising a limit is allowed, but it needs a reason. The API went from 64M to 128M when it started handling per-user encrypted databases and spreadsheet parsing, and the reason sits in a comment next to the number. That is the difference between a budget and a guess.

How I got here

The box didn’t start this small. Last year, before I built the current stack, I ran a nearly static personal site on a machine with 4 GB of RAM: Apache, PHP, WordPress and MySQL. No page cache. Every visit went through PHP and MySQL, for content that almost never changed.

It ran out of memory again and again. I found out through the mail client on my phone, which started throwing errors; the cause was MySQL being killed by the OOM killer. When the site did answer, it sometimes took 10 to 15 seconds.

The mail server kept its virtual mailboxes and authentication in that same MySQL. It was a habit from a previous setup, where editing map files for every new mailbox and group had worn me out. Here I didn’t need it, and it meant that when MySQL died, my mail went with it.

The mistake was mine, and it was specific: a dynamic CMS with no cache, serving pages that don’t change. The obvious fix was more resources. I didn’t take it. I rewrote what had to be rewritten, put whatever could be static behind a CDN with caching, and kept on the server only what has to run there.

The WordPress site itself became a static site. I had used it as a design system, never as a blog, and its content didn’t change, so a PHP stack serving the same HTML every time was overkill, and one more thing to keep updated.

Today the box has a quarter of that RAM. That is where the rule in the previous section comes from.

What lives on the box, and what doesn’t

The rule I ended up with: if it can be a file served from a CDN, it doesn’t get RAM.

Every frontend I run is static and lives on Cloudflare Pages. One of my tools also has a small Worker that parses HTTP headers at the edge, so a request doesn’t make a useless trip to the backend.

What stays on the server is what holds state or has to be reachable on its own ports: the ingress, the SSO, the database, the API, and the mail server, which has to be exposed on its own ports. Mail configuration lives in plain files now; there is no database behind it.

Here is what each one is allowed and what it actually takes, rounded:

ServiceLimitReal use
Ingress (Traefik)none~15 MB
Database (Postgres)none~45 MB
API (Rust)128 MB~5 MB
Mail server512 MB~35 MB
SSO server600 MB~160 MB
SSO worker300 MB~270 MB
Certificate renewalnoneunder 1 MB

The table is a snapshot, not an average. It shows three things I’ll come back to: the limits don’t add up to the RAM of the machine, one row is much closer to its ceiling than the others, and some rows have no ceiling at all.

Where it’s still tight

The SSO worker has a 300 MB ceiling and sits close to it. On 24 September a scheduled metadata import pushed it over. The kernel logged Memory cgroup out of memory and killed the process, the worker shut down, and the container restarted itself. Docker still reported OOMKilled=false and exit code 0, because the kill hit a child process and not the main one, so docker inspect alone told me nothing. The container has restarted nine times since early July; I have only proved the cause of the last one.

Nothing broke, and I found out only because I went looking. That is the kind of failure a single-user box hides. Last year the OOMs were constant and I noticed them from my phone; today they are occasional and invisible. Better, not finished.

I haven’t raised the limit. The worker isn’t the problem so much as the SSO, which is heavier than what I need. I plan to move to a lighter one, as long as security doesn’t suffer.

The ingress and the database have no ceiling at all. That is a choice: I want to see how the stack holds up when they grow, instead of capping them out of habit. The price is that a runaway database has no cgroup of its own to be killed in, so the kernel picks a victim for the whole machine.

The limits add up to about 1.5 GB on a machine with 1 GB. They are ceilings, not reservations, and it works because the services never peak together. That is an assumption, and I’d rather write it down than hide it. There is swap too, as a buffer.

What to take from it

This doesn’t prove that everything fits in 1 GB. I am the only user of most of it, and the SSO load is close to nothing. It shows that a habit holds up.

  • Set a memory ceiling before a service runs, then compare it with what it really uses.
  • Raising a limit is fine. Write the reason next to the number.
  • If it can be a static file on a CDN, it doesn’t need RAM.
  • Docker’s OOMKilled flag only describes the main process. When it says no, read the kernel log.
  • Limits that add up to more than the machine, or a service with no ceiling, are assumptions. Write them down.
  • Generated code runs on the first try, and that is the trap: nobody asks what it needs.
  • Don’t just buy the size an AI agent suggests. Try a small virtual server first: it forces you to optimize and to think about cost before the invoice does.

← all posts