Heartbeat monitoring and uptime monitoring sound like the same thing. They are opposites. An uptime check is a request a monitoring service makes to your website from the outside. It asks "are you responding?" and alerts you when the answer stops coming. A heartbeat check runs the other way. Your own systems, the backup script, the cron job, the queue worker, ping the monitoring service each time they run, and the alert fires when the pings stop.
The direction decides what each one can see. Uptime monitoring sees what your customers see. Heartbeat monitoring sees what nobody sees: the scheduled jobs that run at 3am, produce no page, and can fail for weeks without a visible symptom. This post covers how each works, the failures each catches and misses, and why the honest answer to "which one do you need" is usually both.
Uptime monitoring: the view from outside
An uptime monitor requests a URL on a schedule and judges the response. Did the server answer at all. Did it answer quickly enough. Did it return a healthy status code rather than a 500. Better checks also confirm the page contains what it should, because a 200 serving an error message is still a failure.
Because the probe comes from outside, it exercises the whole chain a visitor depends on: DNS, the network path, the TLS certificate, the web server, the application. If any link breaks, the check fails. That is exactly the property you want. It measures the same journey a customer makes.
What it cannot do is see anything that never touches a page. A large share of what keeps a business running has no URL to probe, and that is where uptime monitoring goes quiet.
The failures uptime checks never see
Consider three jobs. All are common, and all are invisible from the outside.
- The nightly database backup. A disk fills up, or a database password changes, and the backup script starts exiting early. The website is unaffected. Every page loads. Every uptime check passes. As a hypothetical but entirely typical example: the script has been failing silently for six weeks when the server dies, and the newest restorable backup is from before six weeks of orders.
- The certificate renewal cron. Renewal runs on a schedule and fails without symptoms: the timer stopped after a server change, or a DNS challenge broke. Nothing looks wrong until the certificate expires and every visitor meets a full-screen warning. The uptime check only complains on expiry day, which is the day you least wanted to find out.
- The order export job. An hourly job pushes new orders to a warehouse system. After a deploy, the queue worker it depends on never restarts. The shop takes orders all week. Nothing ships. The site was up the entire time.
The pattern is the same each time. The job produces no page, so there is nothing to probe, and when it stops, nothing looks different. Failure produces an absence, and absences do not trigger alerts by themselves.
Heartbeat monitoring: the job reports in
A heartbeat monitor reverses the direction. Instead of the monitoring service calling your systems, your systems call the monitoring service. When you create the monitor you get a secret ping URL. Your job requests that URL every time it completes. A plain GET is enough. In a crontab it is one extra line:
*/5 * * * * curl -fsS --retry 3 https://www.tldtrack.com/ping/YOUR_TOKEN > /dev/null
The flags matter slightly: -f treats an HTTP error as a failure, -sS keeps it quiet except for errors, and --retry 3 absorbs a momentary network blip. For a script rather than a bare cron line, put the ping on the last line. That way the ping means "the job finished", not "the job started".
Two numbers define the monitor. The period is how often a ping should arrive: every five minutes for a queue worker, daily for a backup. The grace is the extra allowance after the period before the alert fires. It absorbs normal jitter, such as a backup that sometimes takes forty minutes instead of twenty. When no ping arrives within the period plus the grace, the monitor goes down and you are alerted. The moment the next ping arrives, it recovers instantly. There is no reset step and no dashboard to visit.
That is the whole mechanism. It is deliberately dumb, which is why it is dependable: the only thing that can satisfy the monitor is the job actually running to completion.
What heartbeats miss
A heartbeat proves the job ran and reached its final line. It does not prove the output was any good. A backup script can complete while writing an empty archive, which is why heartbeats pair with occasional restore tests rather than replace them. If you want more assurance, move the ping behind a verification step, for example after checking the archive is larger than some minimum.
A heartbeat also says nothing about the website itself. A fleet of green heartbeats is entirely compatible with a site that has been down for an hour. And the ping URL is a credential of sorts: anyone who has it can mark the job healthy, so keep it out of public repositories.
Which one do you need?
Both, because they cover different audiences. Uptime monitoring covers what customers see. Heartbeat monitoring covers what nobody sees. A team that runs only uptime checks finds out about backup failures at restore time, which is the most expensive possible moment. A team that runs only heartbeats finds out about outages from customers. The overlap between the two is close to zero, so one cannot substitute for the other.
The good news is that heartbeats are cheap to add. Each one costs a single line in a script you already run. Start with the jobs whose silent failure would hurt most: backups, certificate renewals, exports and anything that moves money or data on a schedule. A perfect 200 from the homepage hides more than most people expect; we catalogued the rest of what it hides in why uptime monitoring isn't enough.
How TLDTrack runs both
TLDTrack runs the two side by side. Uptime checks run every 5 minutes against every monitored site, tightening to every minute while a site is down, and an outage is confirmed from three separate regions before you are alerted, which filters out local network blips. Heartbeat checks are one of the three kinds of service monitor: set a period anywhere from five minutes to daily, add a grace window, and paste the ping URL into the job. Both kinds report through the same alerting channels: email, SMS, push, Slack, Teams and webhooks. Each monitor keeps its own 90-day history, so a backup that misses now and then shows up as a pattern rather than a one-off.
