Skip to content

Domain and DNS Monitoring: What to Watch, How Often, and How to Never Miss an Alert (2026)

Short answer

Domain and DNS monitoring is the automatic, continuous checking of your domain's DNS records, usually every 1 to 15 minutes, so you know before your customers do when a record changes, disappears or stops resolving. A complete setup also watches the things around DNS that break just as often: domain expiry, nameserver delegation, SSL certificates and DNSSEC.

Detecting the problem is only half the job. The other half is making sure the alert wakes the right engineer and turns into a tracked incident. That is the part an incident management platform like ITOC360 handles, from on-call scheduling and escalation to incident response.

Key takeaways

  • Monitor six record types at minimum: A, AAAA, CNAME, MX, NS and TXT (SPF, DKIM, DMARC), plus SOA for zone changes.

  • Check critical records every 1–5 minutes, less critical ones every 15–60 minutes, and expiry dates once a day.

  • Page only on what hurts customers now: resolution failures and unexpected NS or MX changes. Send expiry warnings to a ticket queue, not a phone.

  • Use at least two monitoring locations so a single network blip does not wake anyone.

Why DNS monitoring matters

When DNS fails, everything behind the domain fails with it: the website, the API, login and email. Users rarely see a "DNS error"; they see a site that is simply gone. Three well-documented incidents show how it happens:

Date

What happened

Impact

Lesson

July 14, 2025

A configuration change at Cloudflare withdrew the routes for its 1.1.1.1 public resolver. The faulty setting had sat dormant since June 6.

Resolver unavailable worldwide for 62 minutes (21:52–22:54 UTC). Internal alerts fired 9 minutes after impact began (post-mortem).

A change can wait weeks before it breaks DNS. Only continuous checks catch it.

October 4, 2021

During maintenance, Facebook's network withdrew the BGP routes to its own DNS servers (engineering post).

Facebook, Instagram and WhatsApp unreachable for about six hours.

External monitoring must run outside your own network, or it fails with you.

October 21, 2016

The Mirai botnet hit DNS provider Dyn with a large DDoS attack.

Major sites including Twitter, Netflix and Reddit were unreachable for hours in parts of the US and Europe.

A single DNS provider is a single point of failure. Monitor resolution from several regions.

The pattern is the same each time: DNS breaks, every dependent service breaks with it, and the clock on customer impact starts immediately. Detection time is what you control, and it is the first number in your MTTD, MTTA and MTTR.

How DNS monitoring works

A DNS monitor sends the same DNS query a browser would, on a schedule, and compares the answer with what you expect. Most tools run four kinds of checks:

  1. Resolution check. Does the name resolve at all, and how fast? A timeout or SERVFAIL means users cannot reach you.

  2. Value check. Does the answer match the expected value, such as the IP in an A record or the mail host in an MX record? A silent change is often the first sign of a hijack or a bad deploy.

  3. Authoritative check. The monitor queries your authoritative nameservers directly, not a cached resolver. This catches problems before caches expire and hide them, and confirms every nameserver returns the same answer.

  4. Propagation check. The same query from several regions shows whether a change has reached everyone, and whether one region sees something different.

Each check runs from more than one location. A failure is usually confirmed from two or more locations before it raises an alert, which filters out local network blips.

Two settings decide how fast you find out: the check interval and the record's TTL (time to live). A record with a one-hour TTL can keep serving a wrong value from caches for up to an hour after you fix it, so low TTLs on critical records shorten both detection and recovery.

Which DNS records to monitor

Monitor every record that a customer-facing service depends on. The table lists what each record does, what failure looks like, and how urgently it should reach a person.

Record

What it does

What a failure looks like

Alert as

A / AAAA

Maps the name to an IPv4 / IPv6 address

Site and API unreachable, or traffic sent to the wrong server

Page immediately

CNAME

Points one name at another, often a CDN or SaaS host

Broken subdomain; a dangling CNAME can enable subdomain takeover

Page for production names, ticket for others

MX

Names the servers that receive your email

Inbound email bounces or goes to an attacker's server

Page on any unexpected change

NS

Delegates the zone to its authoritative nameservers

The whole domain stops resolving, or someone else now controls it

Page immediately

TXT (SPF, DKIM, DMARC)

Proves which servers may send mail for you

Your email lands in spam; spoofing becomes easier

Ticket, same business day

SOA

Holds zone metadata, including the serial number

Serial not increasing means updates are not propagating

Ticket

SRV

Locates services such as SIP or LDAP by host and port

VoIP, directory or internal services cannot be found

Page if the service is customer-facing

CAA

Lists which certificate authorities may issue for the domain

Certificate renewals fail, or an unexpected CA is allowed

Ticket

The last column matters as much as the first. Paging on every TXT change trains your team to ignore DNS alerts; see how to reduce alert fatigue.

Beyond records: the rest of domain monitoring

Record checks catch a broken answer. These five checks catch the problems that break DNS from the outside, and together they make up a complete domain monitoring setup.

Domain expiry

An expired domain stops resolving within days, and after the grace period anyone can register it. Monitoring reads the expiry date from WHOIS or RDAP once a day and warns at 30, 14, 7, 3 and 1 days. Registrar renewal emails are not enough: they go to one inbox, and that person may have left.

Nameserver delegation

The NS records at your registrar (the parent zone) must match the NS records in your own zone. A mismatch after a provider migration leaves some resolvers asking servers that no longer answer. Check the delegation from the top-level domain's servers, not only from your zone.

SSL/TLS certificates

A certificate is tied to the domain name. An expired certificate, or one that no longer matches the name, makes browsers block the site even though DNS resolves correctly. Monitor expiry dates and the certificate chain together with DNS.

DNSSEC

DNSSEC signs your records so resolvers can verify they have not been forged. Its signatures expire and its keys rotate; if either step fails, validating resolvers reject your domain entirely. If you enable DNSSEC, monitor signature validity, not just record values.

Lookalike domains

Attackers register names that resemble yours (typos, swapped letters, other TLDs, look-alike Unicode characters) for phishing. Brand-protection services and certificate transparency logs can flag new registrations and new certificates for such names. Treat these alerts as security tickets, not on-call pages.

DNS performance and security monitoring

A record can hold the right value and still hurt users, either because answers are slow or because someone is abusing your DNS. These two checks cover that gap.

Performance: resolution time

Every page load, API call and login waits for DNS before anything else happens, so a slow answer slows everything behind it. Track response time per monitoring location and per authoritative nameserver:

  • Set a baseline per location. Normal response times differ by region; one global threshold produces false alarms.

  • Alert on a sustained rise, not one slow query. A practical starting rule is a warning when the median stays above twice its baseline for 10–15 minutes. Tune it to your traffic.

  • Page only when queries start timing out. Slow answers go to a ticket; failed answers mean users are already affected.

  • Check each nameserver separately. One slow or silent nameserver adds retries for some users while averages still look healthy.

Security: signs of tampering and abuse

DNS is a common target because controlling it means controlling where users and email go. Watch for these signals:

Signal

What it may mean

Route to

Unexpected change to NS, MX or A records

Domain or DNS account hijacking

On-call page + security team

Different answers from different resolvers or regions

Cache poisoning or spoofing

On-call page + security team

Sudden spike in queries or NXDOMAIN responses

DDoS or random-subdomain attack on your nameservers

On-call page if resolution degrades

Many long or unusual TXT queries from inside your network

DNS tunneling used to move data out

Security ticket from your SIEM or resolver logs

The first three come from external DNS monitoring. The last one comes from resolver logs and security tools, which can also send their alerts to ITOC360, for example through the Google Security Command Center integration. Keeping operations and security alerts in one place means a hijack is handled as one incident, not two.

How often to check, and when to wake someone up

Match the check interval to how fast the failure hurts, and match the alert channel to whether someone must act now.

Check

Interval

Confirm from

Route to

Resolution of production names (A/AAAA, CNAME)

1–5 min

2+ locations

On-call page (phone, SMS, push)

Unexpected NS or MX change

5 min

Authoritative servers

On-call page + security channel

TXT, SOA, SRV, CAA values

15–60 min

1 location

Ticket or team chat

Domain expiry

Daily

WHOIS / RDAP

Ticket at 30 days; page at 3 days if still unrenewed

SSL certificate expiry

Daily

1 location

Ticket at 30 days; page at 3 days if still unrenewed

DNSSEC signature validity

1 hour

Validating resolver

On-call page

Lookalike domains

Daily

Feed or CT logs

Security ticket

The expiry rows show the most useful pattern: start quietly and escalate only if nobody acts. A 30-day warning in a ticket queue is easy to handle; a page at 3 days means the quiet warnings failed. Build that step into your escalation policy, and make sure every page follows a real on-call schedule rather than a personal inbox.

DNS monitoring tools that detect the problem

You probably already run a tool that can check DNS. The table lists common ones, the DNS checks they offer, and how their alerts reach ITOC360.

Tool

DNS and domain checks

Connects to ITOC360

Site24x7

DNS server monitor: resolution, record values, response time; SSL and domain expiry

Native

PRTG (Paessler)

DNS sensor: resolution and record values from your own probes

Native

Pingdom

DNS check type for record values; uptime checks that fail on resolution errors

Native

Datadog

Synthetic DNS tests: record values and response time from managed locations

Native (integration)

Prometheus

DNS probes through the blackbox exporter, alerting via Alertmanager

Native (integration)

Grafana

Alert rules on DNS probe data from Prometheus or other sources

Native (integration)

Zabbix

DNS items for resolution and specific record values

Native

AWS CloudWatch

Route 53 health checks and alarms

Native

Azure Monitor

Azure DNS metrics and availability tests

Native

Google Cloud Monitoring

Uptime checks and Cloud DNS metrics

Native

HetrixTools, UptimeRobot

Domain expiry, nameserver and SSL monitoring

Webhook

Cloudflare

Notifications for DNS and zone events

Webhook

Native integrations are listed in ITOC360's App Store description; see all 50+ on the integrations page. Any other tool can create alerts in ITOC360 through a webhook or Zapier's "Send Alert to ITOC360" action.

From DNS alert to resolved incident with ITOC360

A DNS failure rarely sends one alert. Every monitoring location, every dependent check and every tool fires at once, and an email list turns that into dozens of messages nobody owns. ITOC360 turns them into one incident with one owner.

From DNS alert to resolved incident

Related alerts from all your monitors are grouped by AI correlation into a single incident, routed to the on-call engineer for that service, and escalated to the next level if nobody acknowledges in time. Once someone owns it, the team resolves it in Slack or Microsoft Teams with the runbook attached, and ITOC360 drafts the retrospective from the incident timeline. Start free, with voice and SMS alerts included.

DNS monitoring checklist

Copy this list into your runbook and tick off each item for every production domain.

  • Resolution checks for every customer-facing name, every 1–5 minutes, from at least two regions

  • Value checks on A/AAAA, CNAME, MX and NS against the expected answer

  • Authoritative checks against each of your nameservers, not only a public resolver

  • TXT checks for SPF, DKIM and DMARC records

  • Domain expiry monitoring with warnings at 30, 14, 7, 3 and 1 days

  • Auto-renew enabled at the registrar, with a payment method that will not expire first

  • Registrar lock and, for critical domains, registry lock enabled

  • Multi-factor authentication on every registrar and DNS provider account

  • SSL certificate expiry and chain checks for every HTTPS name

  • DNSSEC signature monitoring, if DNSSEC is enabled

  • TTLs of 300 seconds or less on records you may need to change fast

  • Each check routed to an on-call schedule or a ticket queue, never to a personal inbox

  • An escalation step for expiry warnings nobody has acted on

  • A runbook link attached to every DNS alert

Frequently asked questions

What is DNS monitoring?

DNS monitoring is the continuous, automatic checking of a domain's DNS records to confirm they resolve and return the expected values. It alerts your team when a record fails, changes unexpectedly or stops propagating.

What is a domain name monitoring service?

The term covers two things. In operations, it means monitoring your own domain's DNS records, expiry date, nameservers and SSL certificates. In brand protection, it means watching for new registrations of domains that imitate your brand. Most teams need both, from different tools.

Which DNS records should I monitor?

At minimum A/AAAA, CNAME, MX, NS and TXT records for every production domain, plus SOA to confirm zone updates propagate. Add SRV and CAA if your services rely on them.

How often should DNS be checked?

Every 1–5 minutes for records behind customer-facing services, every 15–60 minutes for lower-risk records, and once a day for domain and certificate expiry.

What is the difference between DNS monitoring and uptime monitoring?

Uptime monitoring checks whether a website or API responds. DNS monitoring checks the step before it: whether the name resolves to the right place. A site can pass an uptime check from one region while DNS is broken or hijacked elsewhere.

How do I make sure DNS alerts reach the right person?

Send them to an incident management platform rather than an email list. ITOC360 groups related DNS alerts into one incident, routes it to the on-call engineer by phone, SMS, push, Slack or Microsoft Teams, and escalates if nobody acknowledges.

Sources