Nobody adds a broken link on purpose, and every site accumulates them anyway. Pages get renamed, products get discontinued, external sites die, migrations misfire, and each small change leaves a link somewhere pointing at nothing. Individually a 404 is a shrug; collectively they are rot, and they always surface at the worst moment, in front of a customer who was one click from what they wanted.
This guide covers where broken links actually come from, what they cost beyond mild embarrassment, how to find and fix them properly, and, the part most guides skip, why a one-off cleanup never stays clean and what to do about that.
Where broken links come from
Your own site generates them constantly. A page is deleted or renamed without a redirect; a slug gets tidied in the CMS and every existing link to the old address dies; products and categories churn out from under the pages that mention them; a PDF is reorganised into a new folder. Redesigns and migrations are the mass-casualty event: a new URL structure without a complete redirect map can 404 years of accumulated links in one deploy.
The outside world generates the rest. Link rot is universal: a meaningful fraction of the external pages you link to will simply stop existing over the next few years, through no action of yours. Sites shut down, articles move behind logins, blogs prune archives. An untouched site with zero edits still grows broken links, which is the fact that makes one-off cleanups a treadmill.
What broken links actually cost
Users first. A 404 lands at the exact moment of intent, someone wanted the thing badly enough to click, and a page with several dead links reads as an abandoned site, which quietly reprices everything else you say. When the broken link sits inside a flow, a "start your trial" button, a checkout step, a booking handoff, it is not cosmetic at all: it is a conversion leak that no uptime check will ever notice, the same category of invisible failure as the rest of the up-but-broken family.
Search second. There is no direct "broken links" penalty, but the indirect costs are real. Internal links pointing at 404s waste crawl budget and throw away the ranking equity those links were passing. A migration that 404s old URLs sheds the value those pages accumulated. And redirecting everything to the homepage does not rescue any of it: blanket redirects get treated as soft 404s, which is to say, as 404s with extra steps.
Finding them: crawl, don't spot-check
Clicking around your own site finds nothing; familiarity steers you down the paths that work. Finding broken links means a crawl: start from the homepage, follow every internal link, request every outbound target, and record every response that comes back 404, 410, or an error. Google Search Console's coverage reports show what Googlebot tripped over, useful, but a lagging indicator of what visitors have already been hitting. Server logs show the 404s people actually reached, which is your triage list sorted by real-world pain.
The catch is that a crawl is a snapshot, and the rot resumes immediately, your team keeps editing, the external web keeps dying. A crawl you ran in March answers questions about March. The only version of this that stays true is a crawl on a schedule.
Fixing them properly
- If the page should exist, restore it. The best redirect is not needing one.
- Otherwise, 301 to the closest live equivalent: the replacement product, the successor article, the parent category. Closest genuinely means closest; a redirect that ignores the visitor's intent is a 404 in disguise.
- Never blanket-redirect to the homepage. It rescues no equity, confuses visitors, and search engines treat it as a soft 404 anyway. Where no equivalent exists, an honest 404 is correct.
- Fix the referring link too, wherever you control it. A redirect is a patch over a wound; the internal link that points at the dead address is the wound. Update it.
- For dead outbound links: point at the source's new home if it moved, link an archived copy if it matters, or unlink if it no longer earns its place.
- Keep a helpful 404 page, search box, top destinations, working navigation, to rescue the visits you cannot redirect: the typos, the ancient bookmarks, the links on sites you will never reach.
Stopping them coming back
Process covers the breakage you cause. Make "what happens to the old URL?" a mandatory question in every deletion, rename and migration, answered before publish, not after the traffic drops; for migrations, build the redirect map first and crawl both before and after the switch.
Monitoring covers the breakage you do not cause. External targets die on their own schedule, other teams edit pages you link to, and the CMS does surprising things on quiet Tuesdays. A scheduled crawl with alerts for newly broken links only turns the problem from an annual archaeology project into a short weekly queue, and new-only alerting is what keeps the queue readable instead of becoming a wall of known issues everyone scrolls past.
Where this fits
Broken-link checking is part of TLDTrack's SEO Health Audit: a scheduled crawl of each site that flags broken internal and outbound links alongside the on-page blockers that travel with them, accidental noindex tags, missing titles, the mistakes covered in our WordPress monitoring guide, with alerts on new problems only. The per-page results sit on the visual sitemap, so you can see at a glance which corners of the site are rotting. It is one layer of the stack in the complete guide to monitoring client websites, and one of the most satisfying to automate: link rot never stops, but it can stop being your job to go looking for it.
