Alert Deduplication: How It Works, How to Design Dedup Keys, and Mistakes to Avoid
A disk fills up at 2 a.m. Your monitoring tool checks it every minute, so by 2:30 it has sent 30 alerts about the same full disk. Without deduplication, that's 30 notifications, 30 tickets, or 30 pages. With it, it's one alert with a counter that says 30.
Alert deduplication is the process of recognizing repeated alerts about the same condition and merging them into a single alert or incident. It's the simplest noise-reduction technique there is, and also the one teams get wrong most often, usually because of a badly designed dedup key.
Quick answer: Alert deduplication merges repeat notifications for the same problem into one open alert. Each incoming alert carries a deduplication key (also called an alias, fingerprint, or idempotency key). If an open alert already has that key, the new event updates it instead of creating a new one. The key should identify what is broken and where, but never include values that change on every event, such as timestamps or metric readings.
This guide covers how deduplication works, how it differs from grouping and correlation, how Prometheus, Zabbix, PagerDuty, and Opsgenie implement it (with config examples), and how to design keys that cut noise without hiding real incidents. Deduplication is the second stage of alert management, right after ingestion. We build ITOC360, an incident orchestration platform listed in the Zabbix integration directory and on the Prometheus integrations page. It deduplicates alerts from Zabbix, Prometheus, Datadog, Grafana, Splunk, and other tools, so the examples here reflect patterns we see in real alert streams.
Key facts at a glance
What it does | Merges repeat alerts for the same condition into one open alert |
Also called | De-duplication, dedup, alert aggregation (loosely) |
Core mechanism | A dedup key compared against currently open alerts |
Key names by tool |
|
What it doesn't do | Link different alerts that share a cause. That's correlation |
Biggest risk | A key so broad it hides a second, unrelated incident |
Metric to watch | Alert-to-incident ratio: raw alerts received per incident created |
What Is Alert Deduplication?
Alert deduplication is how an alerting system decides that a new alert is really an old alert it has already seen. Instead of opening a second alert, it updates the first one: it bumps a counter, adds an entry to the timeline, and refreshes the last-seen time.
Monitoring tools are built to repeat themselves. A check that runs every 60 seconds will keep reporting a problem until someone fixes it. That's useful for the monitoring tool, because it confirms the problem is still there. It's terrible for the person on call, because every repeat looks like a new emergency.
Here's what that looks like in practice. The same check fires five times, and all five events share one key:
Time | Event from monitoring | Dedup key | Result |
|---|---|---|---|
02:10 | Disk /var above 90% on db-prod-02 |
| New alert created, on-call paged |
02:11 | Disk /var above 90% on db-prod-02 |
| Count → 2, no page |
02:12 | Disk /var above 90% on db-prod-02 |
| Count → 3, no page |
02:13 | Disk /var above 90% on db-prod-02 |
| Count → 4, no page |
02:19 | Disk /var back below 80% on db-prod-02 |
| Alert resolved |
One alert, one page, one resolution. The history is still there if anyone needs it.
Where Duplicate Alerts Come From
Duplicates rarely come from one bad rule. In the alert streams we see, they usually come from one of these six sources:
Source | What happens | Typical fix |
|---|---|---|
Polling checks | A check runs every minute and reports the same breach each time | Deduplicate on check + resource |
Monitoring retries | The tool resends an alert because it didn't get an acknowledgment | Make the receiving side idempotent |
Two tools watching one thing | Zabbix and Datadog both alert on the same host's disk | Normalize resource names, then deduplicate across sources |
Duplicate integrations | The same monitoring tool is connected twice, often after a migration | Audit integrations; remove the extra one |
Flapping | A metric hovers around its threshold and alternates between problem and OK | Add hysteresis or a recovery delay |
Log bursts | One error is written thousands of times during an outage | Normalize the message, then deduplicate on error type + service |
Why duplicates cost more than attention
The obvious cost is alert fatigue: people stop trusting pages. Google's SRE book sets the bar clearly: a page should be something a human actually needs to act on (Monitoring Distributed Systems). In our own analysis of 1 million alerts, 60 to 80 percent of the alerts fired in a month needed no meaningful human action: they were duplicates, symptoms of a single upstream cause, or problems that resolved on their own within minutes. There are three less obvious costs:
Duplicate incidents. If every alert opens an incident, two responders can start fixing the same problem in two channels, sometimes with conflicting actions.
Distorted metrics. Incident counts, MTTA, and escalation rates all inflate. A team can look worse on paper because of noise, not because of more failures.
Buried severity changes. When a warning turns critical, that change can get lost among dozens of identical warnings.
Research backs up how hard this gets at scale. An empirical study of real-world alert data from large online service systems found that alert storms make failure diagnosis slow and tedious for engineers (IEEE).
Deduplication vs Grouping vs Correlation vs Suppression
These four terms get mixed up constantly, including in vendor docs. They solve different problems and work best in layers.
Technique | What it merges | Example | Question it answers |
|---|---|---|---|
Deduplication | The same alert repeated | 30 "disk full" events from one host → 1 alert | Have I already seen this exact problem? |
Grouping | Similar alerts that share attributes | High CPU on 12 hosts in one cluster → 1 group | Are these the same kind of problem in one place? |
Correlation | Different alerts with a common cause | DB down + API timeouts + failed health checks → 1 incident | Do these alerts come from the same root cause? |
Suppression | Nothing. It drops or delays alerts | Mute alerts during a planned maintenance window | Should anyone see this at all right now? |
The order matters. Deduplication runs first and is exact: same key, same alert. Grouping and correlation run next and are fuzzier. Suppression is a policy decision on top.
A simple way to remember it: deduplication removes echoes, grouping collects siblings, and correlation finds cause and effect. If your team only has deduplication turned on, a single database outage will still page you once per affected service. That's the gap correlation fills. For the broader picture, see what alert noise is and how to eliminate it.
How Alert Deduplication Works
Every deduplication system, no matter what it calls itself, makes the same three decisions.
1. Which key identifies the alert?
The key is a string built from fields in the alert payload. Some tools let the sender set it directly. Others compute it, often as a hash of labels, and call it a fingerprint. Either way, two alerts with the same key are treated as the same problem.
2. How long does a key stay "open"?
There are two common models:
State-based: the key stays open until the alert is resolved or closed. Any repeat before that is a duplicate. This is the model most on-call teams want.
Window-based: the key stays open for a fixed time window, such as 10 minutes or an hour, starting from the first alert. Repeats inside the window are duplicates. After the window, a repeat opens a new alert.
State-based is safer for ongoing outages. Window-based is useful for noisy, self-healing checks where you want a fresh page if the problem comes back an hour later.
3. What happens to a duplicate?
Usually three things. The counter on the open alert goes up, the event is written to the alert's timeline, and nobody is paged again. Good systems also accept a resolve event with the same key, so the monitoring tool can close the alert on its own when the condition clears.
That last part is where many setups break. If the "problem" event and the "recovery" event don't produce the same key, alerts never auto-resolve and the queue fills with stale alerts.
Fixed windows vs sliding windows
Window-based systems come in two flavors, and the difference shows up during long outages.
Window type | How it works | Good for | Risk |
|---|---|---|---|
Fixed | The window starts at the first alert and ends after a set time, no matter what | Getting a reminder page during a long outage | A 3-hour outage with a 1-hour window pages three times |
Sliding | Every duplicate resets the timer | Chatty checks that fire constantly | An outage that never stops firing never pages again |
State-based | No timer; the key stays open until resolved | Most on-call use cases | Forgotten acknowledged alerts swallow repeats |
Normalize before you deduplicate
Deduplication only works if the same thing is always described the same way. That's rarely true out of the box. One tool reports db-prod-02.us-east.internal, another reports DB-PROD-02. One log line says timeout after 30.4s, the next says timeout after 31.2s.
Normalization fixes that before keys are built:
Resource names: lowercase, strip domain suffixes, map aliases (
payments-apiandpayment-service→ one canonical name).Severity: map
critical,P1,disaster, andsev1to one scale.Messages: replace numbers, IDs, UUIDs, and IP addresses with placeholders before using message text in a key.
Environment: make sure every alert carries
prod,staging, ordev.
Most cross-tool duplicate problems are normalization problems in disguise.
Where deduplication sits in the pipeline
Order matters. A clean IT alerting pipeline runs these steps in sequence. It's the same flow we break down in our guide to automated incident management:
Ingest raw events from every monitoring source.
Normalize names, severity, and environment.
Deduplicate repeats of the same condition.
Correlate different alerts that share a cause.
Create or update the incident.
Route and escalate to the on-call owner, using your escalation policy.
Skip step 2 and step 3 quietly fails. Swap steps 3 and 4 and correlation has to wade through repeats it didn't need to see.
How Popular Tools Handle Deduplication
The concept is the same everywhere, but the field names and edge cases differ. Knowing them matters most when alerts pass through two tools, for example Zabbix into an on-call platform, because each layer can deduplicate, or fail to.
Tool | Key field | How it behaves | Detail worth knowing |
|---|---|---|---|
Label set (fingerprint) | Alerts with identical labels are one alert. | Firing alerts re-notify after | |
Trigger + event tags | "Single" generation mode creates one problem per trigger state; "Multiple" creates one per evaluation | With "All problems if tag values match," an OK event closes only problems whose tag matches | |
| Events with the same key update the open alert | After the alert resolves, the same key opens a new alert. Keys only match within the same routing key | |
Opsgenie (shutting down April 2027) |
| One open alert per alias; repeats raise the count | The count and activity log stop at 100 occurrences, though deduplication continues |
Tool behavior as documented in September 2026.
The PagerDuty routing-key rule catches a lot of teams. If the same problem arrives through two integrations, it won't deduplicate, even with identical keys. The Zabbix setting is the opposite trap: in "Multiple" mode, Zabbix itself sends a new problem event on every evaluation, so the tool downstream has to do the deduplication.
How ITOC360 handles deduplication
Most of the pitfalls above come from deduplicating in several tools at once, each with its own rules. ITOC360 takes a different approach: it sits between your monitoring stack and your on-call team, so repeats are handled once, in one place, for every source.
One alert per condition. Repeated triggers for the same problem merge into a single alert instead of paging again.
One incident per cause. Related alerts are then correlated into one incident, so a database outage doesn't page five teams.
Every source, one flow. Alerts from Zabbix, Datadog, Grafana, Prometheus, New Relic, Amazon CloudWatch, and other tools land in the same response flow, so you don't maintain separate logic in each one.
The right person, fast. The incident goes to whoever is on call, by voice call, SMS, push, Slack, Microsoft Teams, or email, and escalates automatically if nobody acknowledges.
As one ITOC360 customer put it in a review: "Before ITOC360, our team was constantly woken up by non-critical notifications. Now, the de-duplication and correlation group them perfectly."
Configuration Examples
These are minimal, copy-ready starting points. Adjust names and labels to your environment.
Prometheus Alertmanager: control identity with labels
In Alertmanager, an alert's identity is its label set. Keep changing values out of labels (put them in annotations instead), and use group_by to batch related alerts into one notification. For broader rule design, the Prometheus project's alerting best practices are worth a read:
route:
receiver: oncall
group_by: [alertname, cluster, service]
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
inhibit_rules:
- source_matchers: [severity="critical"]
target_matchers: [severity="warning"]
equal: [alertname, instance]The inhibit rule mutes the warning version of an alert while the critical version for the same instance is firing. That's suppression, not deduplication, but it removes the most common "two alerts, one problem" case.
PagerDuty Events API v2: same key for trigger and resolve
{
"routing_key": "YOUR_ROUTING_KEY",
"event_action": "trigger",
"dedup_key": "disk-full:/var:db-prod-02:prod",
"payload": {
"summary": "Disk /var above 90% on db-prod-02",
"source": "db-prod-02",
"severity": "critical"
}
}Send the recovery with "event_action": "resolve" and the exact same dedup_key, through the same routing key.
Zabbix: one problem per condition
For most infrastructure triggers, use these settings:
PROBLEM event generation mode: Single
OK event closes: All problems
For log triggers that should track each error separately, switch to Multiple and All problems if tag values match, with a tag such as error_code extracted from the log line. Then make sure the downstream on-call tool deduplicates on that same tag. ITOC360's native Zabbix media type forwards both problem and recovery events to ITOC360 (setup guide).
A simple key builder
If you send alerts through custom webhooks, build the key in one place, with normalization baked in:
import re
def dedup_key(alert: dict) -> str:
check = alert["check"].lower().strip()
resource = alert["host"].lower().split(".")[0] # drop domain suffix
scope = alert.get("mount") or alert.get("service") or "-"
env = alert.get("env", "prod").lower()
return f"{check}:{scope}:{resource}:{env}"
def normalize_message(msg: str) -> str:
msg = re.sub(r"\b[0-9a-f]{8}-[0-9a-f-]{27}\b", "<uuid>", msg, flags=re.I)
msg = re.sub(r"\d+(\.\d+)?", "<n>", msg)
return msgdedup_key({"check": "disk-full", "host": "DB-PROD-02.us-east.internal", "mount": "/var"}) returns disk-full:/var:db-prod-02:prod every time, whichever tool sent it.
How to Design a Good Deduplication Key
A dedup key should answer two questions: what is broken, and where. If two alerts would need different people or different fixes, they need different keys. If they wouldn't, they should share one.
A reliable pattern is:
<check or alert name>:<resource>:<scope>Alert | Bad key | Why it fails | Better key |
|---|---|---|---|
Disk full |
| Timestamp makes every event unique, so nothing deduplicates |
|
High CPU |
| Metric value changes each time |
|
Disk full |
| Too broad: a second host's full disk is hidden inside the first alert |
|
Certificate expiry |
| Fine as is | Same |
Failed login spike |
| Can explode into thousands of keys during an attack |
|
Five rules we give every team:
Never include timestamps, counters, or metric values. They change on every event.
Always include the resource. Host, service, cluster, mount point, or URL. Without it, one alert swallows many.
Include the environment. A staging and a production failure should never share a key.
Use the same key for trigger and resolve. Otherwise, alerts never close on their own.
Keep keys readable.
disk-full:/var:db-prod-02tells the responder something. A hash tells them nothing.
Should severity be part of the key?
This is the question teams argue about most, and both obvious answers are wrong.
Severity in the key: a disk at 85% (warning) and the same disk at 95% (critical) become two separate alerts. You get paged twice for one problem, and the warning stays open after the critical resolves.
Severity out of the key, with no other rule: the critical event is merged into the open warning as a counter bump. Nobody is paged for the escalation.
The better pattern is to keep severity out of the key and add one rule: re-notify when severity goes up. The alert stays one record with one timeline, but an increase from warning to critical triggers a fresh page and, if needed, a new escalation level. A decrease just updates the alert quietly.
The same logic applies to occurrence counts. A condition that repeats twice in an hour is noise. The same condition repeating 200 times in five minutes may mean the failure is spreading. Many teams set a count threshold that raises priority automatically. Tie both rules to your severity levels so everyone knows what a change means.
Common Alert Deduplication Mistakes
Most deduplication problems fall into one of two buckets: the key is too narrow, so nothing merges, or too broad, so real problems disappear. The second is worse, because it fails silently.
Over-deduplication hides new incidents. If the key is just
disk-full, the second server that fills up becomes a counter bump on the first alert. Nobody gets paged for it. Always check that one key can't cover two different fixes.Acknowledged alerts that absorb new failures. In state-based systems, an alert that was acknowledged and forgotten keeps swallowing repeats for days. Set an auto-close policy or a maximum alert age.
Flapping checks. A check that goes problem, OK, problem every few minutes creates a new alert each cycle, because each OK resolves the key. Fix the threshold or add a recovery delay rather than widening the key.
Recovery events with a different key. Common with custom webhooks. The problem arrives as
db-prod-02 disk fulland the recovery asdb-prod-02 disk ok. The alert never resolves.Deduplicating in two places at once. If Alertmanager groups alerts and the on-call tool also deduplicates, a grouped notification can hide which individual alerts resolved. Decide which layer owns deduplication and keep the other simple.
Treating deduplication as the whole solution. Deduplication removes repeats. It does nothing for the 30 different alerts one database outage causes. That's a job for alert routing and correlation.
Migrating Deduplication Rules Off Opsgenie
Opsgenie shuts down on April 5, 2027, and deduplication is one of the easiest things to break during a move. Opsgenie's alias logic is often configured inside integrations and never written down. When teams switch tools, duplicate pages come back overnight.
Before you migrate, work through this checklist:
Export every integration and note how it sets
alias: a fixed value, a payload field, or Opsgenie's default.Find integrations with no alias set. Those alerts were never deduplicated, and the new tool is a chance to fix that.
Map each alias pattern to the deduplication key field in your new platform.
Check whether the new tool is state-based or window-based. Moving from state-based Opsgenie to a window-based tool changes when repeat pages fire.
Confirm recovery events still produce the same key, so alerts auto-close.
Run both tools in parallel for at least a week and compare alert counts per integration.
For the full timeline and a tool-by-tool comparison, see our Opsgenie end-of-life guide and Opsgenie alternatives.
How to Measure Whether Deduplication Is Working
You can't tune what you don't measure. Track these four numbers per integration, weekly at first and monthly once things settle.
Metric | How to calculate it | What good looks like |
|---|---|---|
Deduplication rate | Duplicate events ÷ total events received | High for polling checks, low for event-based sources like logs |
Alert-to-incident ratio | Alerts created ÷ incidents opened | Falling over time |
Pages per incident | Notifications sent ÷ incidents | As close to 1 as your escalation policy allows |
Missed-incident reviews | Incidents found late because an alert was merged into another | Zero. Any case here means a key is too broad |
False grouping rate | Alerts a responder had to split out of a merged alert ÷ merged alerts | Under a few percent, and falling |
Reopen rate | Alerts that resolve and reopen within one hour ÷ resolved alerts | Low. A high rate points to flapping checks |
The last metric is the one most teams forget. A high deduplication rate looks great on a dashboard, but if it hides a second outage, it's doing damage. Add a question to every post-incident review: "Was any alert for this incident merged into something else?"
For reference, across 14 engineering teams on ITOC360 in Q2 2026, correlating alerts before paging cut pages per incident from 8.4 to 1.2. Deduplication alone gets you part of the way. The rest comes from correlating different alerts that share a cause, and from tracking MTTA to confirm responders are acknowledging faster, not just receiving less.
Test deduplication rules before you roll them out
A bad key can hide an outage, so treat key changes like code changes. Before rolling out a new rule, replay last month's alerts through it and check five cases:
Same condition, same resource: merges into one alert.
Same condition, different resource: stays separate.
Same condition, different environment: stays separate.
Recovery event: resolves the open alert automatically.
Severity increase: triggers a new notification.
Start with one noisy integration, compare pages before and after for a week, and only then expand. Keep a short changelog of key changes. When an incident is missed, it's usually the first place to look.
Frequently asked questions
What is alert deduplication?
Alert deduplication merges repeated alerts about the same condition into one open alert. New events with the same deduplication key update the existing alert instead of creating a new notification.
What is a deduplication key?
A deduplication key is a string that identifies an alert, usually built from the check name, the affected resource, and the environment. Tools call it an alias, dedup_key, fingerprint, or idempotency key.
What is the difference between alert deduplication and alert correlation?
Deduplication merges identical alerts. Correlation links different alerts that share a root cause, such as a database outage and the API timeouts it causes.
What is the difference between deduplication and grouping?
Deduplication merges the exact same alert. Grouping collects similar alerts, such as high CPU on several hosts in one cluster, into a single notification or view.
Can deduplication hide real incidents?
Yes. If a key is too broad, a new failure on a different resource can be merged into an existing alert. Include the resource and environment in every key to prevent this.
Should the dedup key include a timestamp?
No. A timestamp makes every event unique, so nothing is ever deduplicated.
What happens to a duplicate alert after the original is resolved?
In most state-based tools, such as PagerDuty and Opsgenie, a new trigger with the same key opens a new alert once the original is resolved.
Should severity be included in the deduplication key?
Usually not. Keep severity out of the key so one problem stays one alert, and add a rule that re-notifies when severity increases.
What is the difference between a fixed and a sliding deduplication window?
A fixed window ends a set time after the first alert, so long outages page again periodically. A sliding window resets with every duplicate, so a problem that keeps firing never pages again until it stops.
Can alerts from two different monitoring tools be deduplicated?
Yes, if both describe the resource and condition the same way. Normalize host names, severity, and environment first, then build the key from the normalized fields.
Final Thoughts
Deduplication is the cheapest noise fix you'll ever make, and one of the easiest to get subtly wrong. Build keys from what's broken and where. Keep timestamps and metric values out. Make sure trigger and recovery events share a key. Then review, at least once a quarter, whether any key has grown broad enough to hide a real incident.
Once repeats are under control, the next step is correlation: turning many different alerts into one incident. That's where on-call load really drops, especially when the result feeds a clear on-call and escalation flow. If you'd like to see how that works across Zabbix, Prometheus, Datadog, and your other tools, try ITOC360 free.
Related Reading
We Processed 1 Million Alerts. Here's What We Learned About Noise
What Is IT Alerting? A Practical Guide for Engineering Teams
Sources
Prometheus, Alertmanager configuration
Prometheus, Alerting best practices
PagerDuty, Events API v2: Send an alert event
PagerDuty, Event management
Atlassian, What is alert de-duplication? (Opsgenie)
Google, Site Reliability Engineering: Monitoring Distributed Systems
IEEE Xplore, Understanding and Handling Alert Storm for Online Service Systems
ITOC360, We processed 1 million alerts: what we learned about noise