The message always arrives the same way: a client, slightly too calm or not calm at all, asking why their website is down. Usually at 5:47pm on a Friday. What separates agencies that keep clients through outages from agencies that lose them is not how often sites go down, it is whether the response is a practised routine or a scramble.
Here is the five-step playbook we use. Steal it, adapt it, write yours down.
Step 1: Confirm it is actually down
A meaningful fraction of "the site is down" reports are local: the client's office wifi, a cached DNS entry, a corporate proxy, an expired session. Before touching anything, establish whether the site is down for everyone or just for them:
- Check from a network that is not yours and not theirs, for example a phone on mobile data.
- Use an external checker such as our free is it down tool, which tests from outside both networks and tells you the status code.
- Note exactly what "down" looks like: timeout, browser error, blank page, 500, security warning. The symptom is your first diagnostic clue, and you will want it for the client update in step 3.
If it is only down for the client, you have a much smaller problem, and you have avoided "fixing" a healthy server.
Step 2: Find the failing layer
Work down the stack in order. Each layer has a 30-second test, and the symptom usually points at the right one first:
| Symptom | Likely layer | First check |
|---|---|---|
| Domain parking page or NXDOMAIN | Domain expired | WHOIS lookup on the domain |
| Timeout, DNS error | DNS | dig example.com, or a propagation check if records changed recently |
| Connection refused or timeout with healthy DNS | Server / hosting | Host status page, can you SSH in? |
| Full-screen browser security warning | SSL certificate | Check the certificate expiry, see our expired certificate guide |
| 500 error or white screen | Application / CMS | Server error logs, was anything deployed or updated today? |
| Painfully slow, intermittent errors | Resources | Disk space, memory, database connections, traffic spike |
Two questions shortcut most diagnoses: what changed recently (deployments, plugin updates, DNS edits, migrations) and when exactly did it stop working. If your monitoring keeps a change history of DNS and content, the answer is often sitting in the alert log already.
Step 3: Tell the client before they chase you
Send the first update as soon as you have confirmed the outage, not when you have fixed it. Waiting until you have good news reads as silence from the client's side, and silence is what they remember.
First update template: "We're aware yoursite.com is currently unreachable and we're investigating now. It appears to be [hosting-level / DNS / certificate] related. Next update within 30 minutes, sooner if it's resolved."
Then keep the cadence you promised, even if the update is "still working on it, here is what we've ruled out". If the client has a revenue-critical site, a branded status page they can check themselves saves everyone the phone calls.
Step 4: Fix the common causes
- Expired domain: renew immediately and expect up to a few hours of DNS re-propagation. Then move the renewal date into a tracked system so it never surprises you again.
- Expired certificate: re-run issuance, confirm the challenge passes, reload the web server, and verify from outside that the new certificate is being served.
- Hosting outage: confirm on the host's status page, escalate a ticket, and spend your energy on client communication; the fix is out of your hands.
- DNS misconfiguration: restore from your recorded baseline. This is the step where teams without a baseline lose hours reconstructing records from memory.
- Broken update or deployment: roll back first, diagnose second. On WordPress, disable the most recently changed plugin or theme via the filesystem if the admin is unreachable.
- Resource exhaustion: clear space or raise limits to restore service, then schedule the real capacity conversation for daylight hours.
Step 5: Close the loop after recovery
Ten minutes of aftercare converts a bad day into renewed trust:
- Send a short plain-language summary: what happened, what you did, what will prevent a repeat.
- Write the incident into a per-site runbook: hosting, registrar, DNS provider, key contacts, and what broke this time. The next engineer to touch this site should not start from zero.
- Close the monitoring gap that let the client find out first, if there was one.
The real fix: find out before the client does
Every step above gets easier when the alert comes from your monitoring instead of your client. Detection in minutes instead of hours, a recorded history of what changed, and expiry warnings weeks ahead mean most of these incidents either never happen or are fixed before anyone notices. That is the case we make in detail in our complete guide to monitoring client websites, and it is what TLDTrack does for agencies out of the box: uptime with sensible retries, SSL and domain expiry alerts, DNS change history and the layers uptime checks miss, across every client site in one dashboard.
