Skip to main content

Incidents

When a failure is significant enough to track as more than a single alert, it shows up here as an incident, usually reached from Overview or an alert in Alerts & Notifications.

What you'll see

Each incident has a severity (SEV-1 through SEV-3), a status (Investigating, Monitoring, or Resolved), a summary, when it started, how many events have failed because of it, and which services are affected.

Investigating

  • Root cause: the traces most likely responsible, with a link straight into the Trace Graph for each.
  • Propagation timeline: how the incident spread across services over time.

Once you've fixed the underlying issue, an admin/operator can resolve the incident.

Replay

Coming soon Replaying dead-lettered events straight from the dashboard isn't available yet. For now, use the DLQ Center to see what failed and why, then fix the underlying issue in your consumer or producer and let it reprocess normally.