Skip to content

Alert Deduplication: How It Works, How to Design Dedup Keys, and Mistakes to Avoid

A disk fills up at 2 a.m. Your monitoring tool checks it every minute, so by 2:30 it has sent 30 alerts about the same full disk. Without deduplication, that's 30 notifications, 30 tickets, or 30 pages. With it, it's one alert with a counter that says 30.

Alert deduplication is the process of recognizing repeated alerts about the same condition and merging them into a single alert or incident. It's the simplest noise-reduction technique there is, and also the one teams get wrong most often, usually because of a badly designed dedup key.

Quick answer: Alert deduplication merges repeat notifications for the same problem into one open alert. Each incoming alert carries a deduplication key (also called an alias, fingerprint, or idempotency key). If an open alert already has that key, the new event updates it instead of creating a new one. The key should identify what is broken and where, but never include values that change on every event, such as timestamps or metric readings.

This guide covers how deduplication works, how it differs from grouping and correlation, how Prometheus, Zabbix, PagerDuty, and Opsgenie implement it (with config examples), and how to design keys that cut noise without hiding real incidents. Deduplication is the second stage of alert management, right after ingestion. We build ITOC360, an incident orchestration platform listed in the Zabbix integration directory and on the Prometheus integrations page. It deduplicates alerts from Zabbix, Prometheus, Datadog, Grafana, Splunk, and other tools, so the examples here reflect patterns we see in real alert streams.

Key facts at a glance

What it does

Merges repeat alerts for the same condition into one open alert

Also called

De-duplication, dedup, alert aggregation (loosely)

Core mechanism

A dedup key compared against currently open alerts

Key names by tool

alias (Opsgenie, Jira Service Management), dedup_key (PagerDuty), fingerprint (Prometheus Alertmanager)

What it doesn't do

Link different alerts that share a cause. That's correlation

Biggest risk

A key so broad it hides a second, unrelated incident

Metric to watch

Alert-to-incident ratio: raw alerts received per incident created

What Is Alert Deduplication?

Alert deduplication is how an alerting system decides that a new alert is really an old alert it has already seen. Instead of opening a second alert, it updates the first one: it bumps a counter, adds an entry to the timeline, and refreshes the last-seen time.

Monitoring tools are built to repeat themselves. A check that runs every 60 seconds will keep reporting a problem until someone fixes it. That's useful for the monitoring tool, because it confirms the problem is still there. It's terrible for the person on call, because every repeat looks like a new emergency.

Here's what that looks like in practice. The same check fires five times, and all five events share one key:

Time

Event from monitoring

Dedup key

Result

02:10

Disk /var above 90% on db-prod-02

disk-full:/var:db-prod-02

New alert created, on-call paged

02:11

Disk /var above 90% on db-prod-02

disk-full:/var:db-prod-02

Count → 2, no page

02:12

Disk /var above 90% on db-prod-02

disk-full:/var:db-prod-02

Count → 3, no page

02:13

Disk /var above 90% on db-prod-02

disk-full:/var:db-prod-02

Count → 4, no page

02:19

Disk /var back below 80% on db-prod-02

disk-full:/var:db-prod-02

Alert resolved

One alert, one page, one resolution. The history is still there if anyone needs it.

Where Duplicate Alerts Come From

Duplicates rarely come from one bad rule. In the alert streams we see, they usually come from one of these six sources:

Source

What happens

Typical fix

Polling checks

A check runs every minute and reports the same breach each time

Deduplicate on check + resource

Monitoring retries

The tool resends an alert because it didn't get an acknowledgment

Make the receiving side idempotent

Two tools watching one thing

Zabbix and Datadog both alert on the same host's disk

Normalize resource names, then deduplicate across sources

Duplicate integrations

The same monitoring tool is connected twice, often after a migration

Audit integrations; remove the extra one

Flapping

A metric hovers around its threshold and alternates between problem and OK

Add hysteresis or a recovery delay

Log bursts

One error is written thousands of times during an outage

Normalize the message, then deduplicate on error type + service

Why duplicates cost more than attention

The obvious cost is alert fatigue: people stop trusting pages. Google's SRE book sets the bar clearly: a page should be something a human actually needs to act on (Monitoring Distributed Systems). In our own analysis of 1 million alerts, 60 to 80 percent of the alerts fired in a month needed no meaningful human action: they were duplicates, symptoms of a single upstream cause, or problems that resolved on their own within minutes. There are three less obvious costs:

  • Duplicate incidents. If every alert opens an incident, two responders can start fixing the same problem in two channels, sometimes with conflicting actions.

  • Distorted metrics. Incident counts, MTTA, and escalation rates all inflate. A team can look worse on paper because of noise, not because of more failures.

  • Buried severity changes. When a warning turns critical, that change can get lost among dozens of identical warnings.

Research backs up how hard this gets at scale. An empirical study of real-world alert data from large online service systems found that alert storms make failure diagnosis slow and tedious for engineers (IEEE).

Deduplication vs Grouping vs Correlation vs Suppression

These four terms get mixed up constantly, including in vendor docs. They solve different problems and work best in layers.

Technique

What it merges

Example

Question it answers

Deduplication

The same alert repeated

30 "disk full" events from one host → 1 alert

Have I already seen this exact problem?

Grouping

Similar alerts that share attributes

High CPU on 12 hosts in one cluster → 1 group

Are these the same kind of problem in one place?

Correlation

Different alerts with a common cause

DB down + API timeouts + failed health checks → 1 incident

Do these alerts come from the same root cause?

Suppression

Nothing. It drops or delays alerts

Mute alerts during a planned maintenance window

Should anyone see this at all right now?

The order matters. Deduplication runs first and is exact: same key, same alert. Grouping and correlation run next and are fuzzier. Suppression is a policy decision on top.

A simple way to remember it: deduplication removes echoes, grouping collects siblings, and correlation finds cause and effect. If your team only has deduplication turned on, a single database outage will still page you once per affected service. That's the gap correlation fills. For the broader picture, see what alert noise is and how to eliminate it.

How Alert Deduplication Works

Every deduplication system, no matter what it calls itself, makes the same three decisions.

1. Which key identifies the alert?

The key is a string built from fields in the alert payload. Some tools let the sender set it directly. Others compute it, often as a hash of labels, and call it a fingerprint. Either way, two alerts with the same key are treated as the same problem.

2. How long does a key stay "open"?

There are two common models:

  • State-based: the key stays open until the alert is resolved or closed. Any repeat before that is a duplicate. This is the model most on-call teams want.

  • Window-based: the key stays open for a fixed time window, such as 10 minutes or an hour, starting from the first alert. Repeats inside the window are duplicates. After the window, a repeat opens a new alert.

State-based is safer for ongoing outages. Window-based is useful for noisy, self-healing checks where you want a fresh page if the problem comes back an hour later.

3. What happens to a duplicate?

Usually three things. The counter on the open alert goes up, the event is written to the alert's timeline, and nobody is paged again. Good systems also accept a resolve event with the same key, so the monitoring tool can close the alert on its own when the condition clears.

That last part is where many setups break. If the "problem" event and the "recovery" event don't produce the same key, alerts never auto-resolve and the queue fills with stale alerts.

Fixed windows vs sliding windows

Window-based systems come in two flavors, and the difference shows up during long outages.

Window type

How it works

Good for

Risk

Fixed

The window starts at the first alert and ends after a set time, no matter what

Getting a reminder page during a long outage

A 3-hour outage with a 1-hour window pages three times

Sliding

Every duplicate resets the timer

Chatty checks that fire constantly

An outage that never stops firing never pages again

State-based

No timer; the key stays open until resolved

Most on-call use cases

Forgotten acknowledged alerts swallow repeats

Normalize before you deduplicate

Deduplication only works if the same thing is always described the same way. That's rarely true out of the box. One tool reports db-prod-02.us-east.internal, another reports DB-PROD-02. One log line says timeout after 30.4s, the next says timeout after 31.2s.

Normalization fixes that before keys are built:

  • Resource names: lowercase, strip domain suffixes, map aliases (payments-api and payment-service → one canonical name).

  • Severity: map critical, P1, disaster, and sev1 to one scale.

  • Messages: replace numbers, IDs, UUIDs, and IP addresses with placeholders before using message text in a key.

  • Environment: make sure every alert carries prod, staging, or dev.

Most cross-tool duplicate problems are normalization problems in disguise.

Where deduplication sits in the pipeline

Order matters. A clean IT alerting pipeline runs these steps in sequence. It's the same flow we break down in our guide to automated incident management:

  1. Ingest raw events from every monitoring source.

  2. Normalize names, severity, and environment.

  3. Deduplicate repeats of the same condition.

  4. Correlate different alerts that share a cause.

  5. Create or update the incident.

  6. Route and escalate to the on-call owner, using your escalation policy.

Skip step 2 and step 3 quietly fails. Swap steps 3 and 4 and correlation has to wade through repeats it didn't need to see.

The concept is the same everywhere, but the field names and edge cases differ. Knowing them matters most when alerts pass through two tools, for example Zabbix into an on-call platform, because each layer can deduplicate, or fail to.

Tool

Key field

How it behaves

Detail worth knowing

Prometheus Alertmanager

Label set (fingerprint)

Alerts with identical labels are one alert. group_by batches alerts into notifications

Firing alerts re-notify after repeat_interval, 4 hours by default

Zabbix

Trigger + event tags

"Single" generation mode creates one problem per trigger state; "Multiple" creates one per evaluation

With "All problems if tag values match," an OK event closes only problems whose tag matches

PagerDuty

dedup_key

Events with the same key update the open alert

After the alert resolves, the same key opens a new alert. Keys only match within the same routing key

Opsgenie (shutting down April 2027)

alias

One open alert per alias; repeats raise the count

The count and activity log stop at 100 occurrences, though deduplication continues

Tool behavior as documented in September 2026.

The PagerDuty routing-key rule catches a lot of teams. If the same problem arrives through two integrations, it won't deduplicate, even with identical keys. The Zabbix setting is the opposite trap: in "Multiple" mode, Zabbix itself sends a new problem event on every evaluation, so the tool downstream has to do the deduplication.

How ITOC360 handles deduplication

Most of the pitfalls above come from deduplicating in several tools at once, each with its own rules. ITOC360 takes a different approach: it sits between your monitoring stack and your on-call team, so repeats are handled once, in one place, for every source.

  • One alert per condition. Repeated triggers for the same problem merge into a single alert instead of paging again.

  • One incident per cause. Related alerts are then correlated into one incident, so a database outage doesn't page five teams.

  • Every source, one flow. Alerts from Zabbix, Datadog, Grafana, Prometheus, New Relic, Amazon CloudWatch, and other tools land in the same response flow, so you don't maintain separate logic in each one.

  • The right person, fast. The incident goes to whoever is on call, by voice call, SMS, push, Slack, Microsoft Teams, or email, and escalates automatically if nobody acknowledges.

As one ITOC360 customer put it in a review: "Before ITOC360, our team was constantly woken up by non-critical notifications. Now, the de-duplication and correlation group them perfectly."

Configuration Examples

These are minimal, copy-ready starting points. Adjust names and labels to your environment.

Prometheus Alertmanager: control identity with labels

In Alertmanager, an alert's identity is its label set. Keep changing values out of labels (put them in annotations instead), and use group_by to batch related alerts into one notification. For broader rule design, the Prometheus project's alerting best practices are worth a read:

route:
  receiver: oncall
  group_by: [alertname, cluster, service]
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h

inhibit_rules:
  - source_matchers: [severity="critical"]
    target_matchers: [severity="warning"]
    equal: [alertname, instance]

The inhibit rule mutes the warning version of an alert while the critical version for the same instance is firing. That's suppression, not deduplication, but it removes the most common "two alerts, one problem" case.

PagerDuty Events API v2: same key for trigger and resolve

{
  "routing_key": "YOUR_ROUTING_KEY",
  "event_action": "trigger",
  "dedup_key": "disk-full:/var:db-prod-02:prod",
  "payload": {
    "summary": "Disk /var above 90% on db-prod-02",
    "source": "db-prod-02",
    "severity": "critical"
  }
}

Send the recovery with "event_action": "resolve" and the exact same dedup_key, through the same routing key.

Zabbix: one problem per condition

For most infrastructure triggers, use these settings:

  • PROBLEM event generation mode: Single

  • OK event closes: All problems

For log triggers that should track each error separately, switch to Multiple and All problems if tag values match, with a tag such as error_code extracted from the log line. Then make sure the downstream on-call tool deduplicates on that same tag. ITOC360's native Zabbix media type forwards both problem and recovery events to ITOC360 (setup guide).

A simple key builder

If you send alerts through custom webhooks, build the key in one place, with normalization baked in:

import re

def dedup_key(alert: dict) -> str:
    check = alert["check"].lower().strip()
    resource = alert["host"].lower().split(".")[0]   # drop domain suffix
    scope = alert.get("mount") or alert.get("service") or "-"
    env = alert.get("env", "prod").lower()
    return f"{check}:{scope}:{resource}:{env}"

def normalize_message(msg: str) -> str:
    msg = re.sub(r"\b[0-9a-f]{8}-[0-9a-f-]{27}\b", "<uuid>", msg, flags=re.I)
    msg = re.sub(r"\d+(\.\d+)?", "<n>", msg)
    return msg

dedup_key({"check": "disk-full", "host": "DB-PROD-02.us-east.internal", "mount": "/var"}) returns disk-full:/var:db-prod-02:prod every time, whichever tool sent it.

How to Design a Good Deduplication Key

A dedup key should answer two questions: what is broken, and where. If two alerts would need different people or different fixes, they need different keys. If they wouldn't, they should share one.

A reliable pattern is:

<check or alert name>:<resource>:<scope>

Alert

Bad key

Why it fails

Better key

Disk full

disk-full-2026-09-25T02:10:14Z

Timestamp makes every event unique, so nothing deduplicates

disk-full:/var:db-prod-02

High CPU

cpu-high-97%

Metric value changes each time

cpu-high:api-prod-07

Disk full

disk-full

Too broad: a second host's full disk is hidden inside the first alert

disk-full:/var:db-prod-02

Certificate expiry

ssl-expiry:checkout.example.com:prod

Fine as is

Same

Failed login spike

auth-fail:{{source_ip}}

Can explode into thousands of keys during an attack

auth-fail-spike:login-service:prod

Five rules we give every team:

  1. Never include timestamps, counters, or metric values. They change on every event.

  2. Always include the resource. Host, service, cluster, mount point, or URL. Without it, one alert swallows many.

  3. Include the environment. A staging and a production failure should never share a key.

  4. Use the same key for trigger and resolve. Otherwise, alerts never close on their own.

  5. Keep keys readable. disk-full:/var:db-prod-02 tells the responder something. A hash tells them nothing.

Should severity be part of the key?

This is the question teams argue about most, and both obvious answers are wrong.

  • Severity in the key: a disk at 85% (warning) and the same disk at 95% (critical) become two separate alerts. You get paged twice for one problem, and the warning stays open after the critical resolves.

  • Severity out of the key, with no other rule: the critical event is merged into the open warning as a counter bump. Nobody is paged for the escalation.

The better pattern is to keep severity out of the key and add one rule: re-notify when severity goes up. The alert stays one record with one timeline, but an increase from warning to critical triggers a fresh page and, if needed, a new escalation level. A decrease just updates the alert quietly.

The same logic applies to occurrence counts. A condition that repeats twice in an hour is noise. The same condition repeating 200 times in five minutes may mean the failure is spreading. Many teams set a count threshold that raises priority automatically. Tie both rules to your severity levels so everyone knows what a change means.

Common Alert Deduplication Mistakes

Most deduplication problems fall into one of two buckets: the key is too narrow, so nothing merges, or too broad, so real problems disappear. The second is worse, because it fails silently.

  • Over-deduplication hides new incidents. If the key is just disk-full, the second server that fills up becomes a counter bump on the first alert. Nobody gets paged for it. Always check that one key can't cover two different fixes.

  • Acknowledged alerts that absorb new failures. In state-based systems, an alert that was acknowledged and forgotten keeps swallowing repeats for days. Set an auto-close policy or a maximum alert age.

  • Flapping checks. A check that goes problem, OK, problem every few minutes creates a new alert each cycle, because each OK resolves the key. Fix the threshold or add a recovery delay rather than widening the key.

  • Recovery events with a different key. Common with custom webhooks. The problem arrives as db-prod-02 disk full and the recovery as db-prod-02 disk ok. The alert never resolves.

  • Deduplicating in two places at once. If Alertmanager groups alerts and the on-call tool also deduplicates, a grouped notification can hide which individual alerts resolved. Decide which layer owns deduplication and keep the other simple.

  • Treating deduplication as the whole solution. Deduplication removes repeats. It does nothing for the 30 different alerts one database outage causes. That's a job for alert routing and correlation.

Migrating Deduplication Rules Off Opsgenie

Opsgenie shuts down on April 5, 2027, and deduplication is one of the easiest things to break during a move. Opsgenie's alias logic is often configured inside integrations and never written down. When teams switch tools, duplicate pages come back overnight.

Before you migrate, work through this checklist:

  • Export every integration and note how it sets alias: a fixed value, a payload field, or Opsgenie's default.

  • Find integrations with no alias set. Those alerts were never deduplicated, and the new tool is a chance to fix that.

  • Map each alias pattern to the deduplication key field in your new platform.

  • Check whether the new tool is state-based or window-based. Moving from state-based Opsgenie to a window-based tool changes when repeat pages fire.

  • Confirm recovery events still produce the same key, so alerts auto-close.

  • Run both tools in parallel for at least a week and compare alert counts per integration.

For the full timeline and a tool-by-tool comparison, see our Opsgenie end-of-life guide and Opsgenie alternatives.

How to Measure Whether Deduplication Is Working

You can't tune what you don't measure. Track these four numbers per integration, weekly at first and monthly once things settle.

Metric

How to calculate it

What good looks like

Deduplication rate

Duplicate events ÷ total events received

High for polling checks, low for event-based sources like logs

Alert-to-incident ratio

Alerts created ÷ incidents opened

Falling over time

Pages per incident

Notifications sent ÷ incidents

As close to 1 as your escalation policy allows

Missed-incident reviews

Incidents found late because an alert was merged into another

Zero. Any case here means a key is too broad

False grouping rate

Alerts a responder had to split out of a merged alert ÷ merged alerts

Under a few percent, and falling

Reopen rate

Alerts that resolve and reopen within one hour ÷ resolved alerts

Low. A high rate points to flapping checks

The last metric is the one most teams forget. A high deduplication rate looks great on a dashboard, but if it hides a second outage, it's doing damage. Add a question to every post-incident review: "Was any alert for this incident merged into something else?"

For reference, across 14 engineering teams on ITOC360 in Q2 2026, correlating alerts before paging cut pages per incident from 8.4 to 1.2. Deduplication alone gets you part of the way. The rest comes from correlating different alerts that share a cause, and from tracking MTTA to confirm responders are acknowledging faster, not just receiving less.

Test deduplication rules before you roll them out

A bad key can hide an outage, so treat key changes like code changes. Before rolling out a new rule, replay last month's alerts through it and check five cases:

  • Same condition, same resource: merges into one alert.

  • Same condition, different resource: stays separate.

  • Same condition, different environment: stays separate.

  • Recovery event: resolves the open alert automatically.

  • Severity increase: triggers a new notification.

Start with one noisy integration, compare pages before and after for a week, and only then expand. Keep a short changelog of key changes. When an incident is missed, it's usually the first place to look.

Frequently asked questions

What is alert deduplication?

Alert deduplication merges repeated alerts about the same condition into one open alert. New events with the same deduplication key update the existing alert instead of creating a new notification.

What is a deduplication key?

A deduplication key is a string that identifies an alert, usually built from the check name, the affected resource, and the environment. Tools call it an alias, dedup_key, fingerprint, or idempotency key.

What is the difference between alert deduplication and alert correlation?

Deduplication merges identical alerts. Correlation links different alerts that share a root cause, such as a database outage and the API timeouts it causes.

What is the difference between deduplication and grouping?

Deduplication merges the exact same alert. Grouping collects similar alerts, such as high CPU on several hosts in one cluster, into a single notification or view.

Can deduplication hide real incidents?

Yes. If a key is too broad, a new failure on a different resource can be merged into an existing alert. Include the resource and environment in every key to prevent this.

Should the dedup key include a timestamp?

No. A timestamp makes every event unique, so nothing is ever deduplicated.

What happens to a duplicate alert after the original is resolved?

In most state-based tools, such as PagerDuty and Opsgenie, a new trigger with the same key opens a new alert once the original is resolved.

Should severity be included in the deduplication key?

Usually not. Keep severity out of the key so one problem stays one alert, and add a rule that re-notifies when severity increases.

What is the difference between a fixed and a sliding deduplication window?

A fixed window ends a set time after the first alert, so long outages page again periodically. A sliding window resets with every duplicate, so a problem that keeps firing never pages again until it stops.

Can alerts from two different monitoring tools be deduplicated?

Yes, if both describe the resource and condition the same way. Normalize host names, severity, and environment first, then build the key from the normalized fields.

Final Thoughts

Deduplication is the cheapest noise fix you'll ever make, and one of the easiest to get subtly wrong. Build keys from what's broken and where. Keep timestamps and metric values out. Make sure trigger and recovery events share a key. Then review, at least once a quarter, whether any key has grown broad enough to hide a real incident.

Once repeats are under control, the next step is correlation: turning many different alerts into one incident. That's where on-call load really drops, especially when the result feeds a clear on-call and escalation flow. If you'd like to see how that works across Zabbix, Prometheus, Datadog, and your other tools, try ITOC360 free.

Sources