Caddy in Docker, and the www certificate that wouldn't issue
I moved this site's web server off the host and into a container, so I can experiment on
the VPS without risking the site. Then I put it on a domain with HTTPS, and the apex got a
certificate in seconds while www kept failing.
Why move it at all
The VPS has two jobs: it serves this site, and it's a machine I want to try things on.
With Caddy installed as a host package, those jobs overlap. The config lives in
/etc/caddy, the content in /var/www/html, and the service is a
systemd unit, so package upgrades, a second web server, firewall changes or Docker
networking tests all touch the same system the site depends on.
The goal was for the site to depend only on one container and a few folders, so that anything else on the host can break without taking it down.
Starting point
| Part | Before |
|---|---|
| Host | Debian 13 VPS, SSH keys only, ufw allowing 22, 80, 443 |
| Web server | Caddy from the official apt repository, run by systemd |
| Content | /var/www/html |
| Exposure | Plain HTTP on port 80, reached by IP. No domain yet. |
Target layout
scp into ~/caddy/site. The container picks it up without a restart.Steps
1. Install Docker
curl -fsSL https://get.docker.com -o get-docker.sh
sudo sh get-docker.sh
sudo usermod -aG docker $USER # then log out and back in
docker run hello-world
Adding my user to the docker group means I don't need sudo for every
command, but it's effectively root access to the host, since anyone in that group can start a
container that mounts /. That's acceptable on a single-admin VPS. I'd think twice
about it on a shared one.
2. Free ports 80 and 443
sudo systemctl stop caddy
sudo systemctl disable caddy
sudo ss -tlnp | grep -E ':80|:443' # no output = ports are free
Only one process can listen on a port, so the host's Caddy had to stop before the container could start. That leaves a gap of a few seconds with nothing serving the site. Here it didn't matter because nothing was live yet. On a real service I'd start the container on a spare port first, test it, and only then switch.
3. Keep config and content on the host
mkdir -p ~/caddy/site
First version of ~/caddy/Caddyfile, still serving HTTP by IP:
:80 {
root * /srv
file_server
}
/srv is a path inside the container. The next step maps the host folder onto it.
4. Run the container
docker run -d \
--name caddy \
--restart unless-stopped \
-p 80:80 \
-p 443:443 \
-v ~/caddy/Caddyfile:/etc/caddy/Caddyfile \
-v ~/caddy/site:/srv \
-v caddy_data:/data \
caddy:latest
| Flag | Why |
|---|---|
--restart unless-stopped | Comes back after a crash or a reboot. Stays down only if I stop it myself. |
-p 80:80 -p 443:443 | Host port → container port. 443 was added in step 5, when HTTPS came in. |
-v ~/caddy/Caddyfile:… | Config is a file on the host that I edit with nano. It isn't baked into the container. |
-v ~/caddy/site:/srv | Content. Deploying means copying files here, with no restart and no rebuild. |
-v caddy_data:/data | Named volume for certificates and the ACME account. Without it, every docker rm throws them away, and re-requesting them repeatedly runs into Let's Encrypt's rate limits. |
Deploying the site itself, from the laptop:
# run on the laptop, where the files are
scp -r site/* user@my-vps:~/caddy/site/
5. Domain and HTTPS
At the registrar, I replaced the default ALIAS record on the apex (it pointed at
the registrar's parking page) with an A record pointing at the VPS. I also checked
it with dig +short nikitchenko.com instead of trusting the browser, which kept
showing the registrar's proxy error for a few minutes from cache.
Then I swapped the Caddyfile's :80 for real names, which is all Caddy needs to request and renew certificates on its own:
nikitchenko.com, www.nikitchenko.com {
root * /srv
file_server
}
I recreated the container with port 443 and the caddy_data volume (the command in step 4).
Symptom
nikitchenko.com got a certificate within seconds. www.nikitchenko.com
failed both challenge types Caddy tried (log trimmed):
challenge failed identifier=www.nikitchenko.com challenge_type=tls-alpn-01
Cannot negotiate ALPN protocol "acme-tls/1" for tls-alpn-01 challenge
challenge failed identifier=www.nikitchenko.com challenge_type=http-01
207.207.210.126: Fetching https://nikitchenko-com.l.ink/.well-known/acme-challenge/…:
Error getting validation data
Hypotheses I checked
| Hypothesis | How I checked | Result |
|---|---|---|
| Port 443 isn't reachable | The apex validated over tls-alpn-01 on the same port and the same container | Ruled out |
www is missing from the Caddyfile | Startup log lists both names under "automatic TLS certificate management" | Ruled out |
www resolves to some other server | The http-01 error names the host Let's Encrypt actually reached: the registrar's l.ink service at 207.207.210.x. That isn't my server. | Confirmed |
Root cause
A new domain at my registrar comes with default records: an ALIAS on the apex
and a wildcard CNAME *.nikitchenko.com → uixie.porkbun.com, both pointing
at the registrar's parking/link-in-bio service. I replaced the apex record and left the
wildcard alone.
A wildcard matches any name that has no record of its own, and that includes www.
So Let's Encrypt, which validates by connecting to wherever the name resolves, landed on the
registrar's servers. Those servers don't speak acme-tls/1 and have never heard of
Caddy's challenge token, so both challenges failed. Caddy and Docker were working correctly.
The DNS record for that one name was wrong.
Fix
One record at the registrar:
A www.nikitchenko.com → (VPS address)
A record for a specific name takes priority over a wildcard, so www now resolves
to the VPS. I didn't have to do anything on the server, because Caddy retries failed issuance with
backoff (60 s, then 120 s). On the retry it first ran the challenge against Let's Encrypt's
staging environment, and only after that succeeded did it request the real
certificate:
obtaining certificate identifier=www.nikitchenko.com ca=acme-staging-v02…
authorization finalized authz_status=valid ca=acme-staging-v02…
authorization finalized authz_status=valid ca=acme-v02…
certificate obtained successfully identifier=www.nikitchenko.com
That staging step keeps a broken setup from burning through production rate limits. I also
ran docker restart caddy "to force a retry", but the logs showed the certificate
had already been issued about two minutes before I did.
Result
| Part | Before | After |
|---|---|---|
| Web server | Host package + systemd unit | One container, restarts itself |
| Config | /etc/caddy/Caddyfile | ~/caddy/Caddyfile, mounted in |
| Content | /var/www/html | ~/caddy/site. Deploy = scp, no restart |
| TLS | None, HTTP by IP | Let's Encrypt for apex and www, renewed automatically, stored in a volume |
Smaller things that went wrong
-
scpfrom the wrong side. My first deploy ran inside the SSH session, so~/Downloadsmeant the server's home folder and the error wasstat local "/home/…/Downloads/…": No such file or directory.scphas to run on the machine that has the files. -
Trusting the browser about DNS. Right after the DNS change the browser still
showed the registrar's proxy error.
dig +shortalready returned the right address, so the browser was serving a cached answer.
What I learned
- Keep the container and its state apart. Config, content and certificates live outside the container, and that's why
docker rmwas safe to run. - Fixing DNS for the apex doesn't fix it for the domain. Check every name you serve, and look for the registrar's default records, especially wildcards.
- ACME errors are specific. This one named the exact host Let's Encrypt reached, which pointed straight at DNS.
- Read the logs before stepping in. The retry had already succeeded by the time I restarted the container.
Loose ends
- Move the
docker runcommand into acompose.yaml, so the setup lives in a file under version control and not in shell history. - Publish
443/udpas well. Caddy enables HTTP/3, but only TCP 443 is published, so browsers fall back to HTTP/2. - Ports published by Docker get their own iptables rules, and those rules bypass ufw. That's fine for 80 and 443, which are meant to be public. Anything published later needs binding to
127.0.0.1or rules in theDOCKER-USERchain. - Remove the disabled host Caddy package, so there's only one web server on the machine.
- Run
caddy fmt. The logs warn that the Caddyfile isn't formatted. - Add uptime monitoring and a public status page.