Incident Management Integration
AxonOps provides several integrations for notifications.
The functionality is accessible via Settings > Integrations
The current integrations are:
- SMTP / Email
- PagerDuty
- Slack
- Microsoft Teams
- ServiceNow
- OpsGenie (coming soon)
- Generic webhooks (coming soon)
- Log file (configurable through
axon-server.yml)

Incident Management Integration
Section titled “Incident Management Integration”AxonOps is designed as a monitoring and alerting system that:
- Detects issues
- Triggers alerts
- Sends recovery events when conditions return to normal
However, AxonOps is not intended to replace dedicated incident management platforms like PagerDuty or OpsGenie.
Incident management platforms provide capabilities such as:
- Converting alerts into incidents with defined workflows
- Escalation policies when initial responders don’t acknowledge
- Repeat notifications until someone takes action
- Acknowledgment to pause notifications while investigating
- Auto-resolution when recovery events arrive
Reducing Alert Fatigue
Section titled “Reducing Alert Fatigue”One of the most valuable features of incident management platforms is alert grouping. When a systemic issue affects your Cassandra or Kafka cluster, it often triggers alerts from multiple nodes simultaneously. Without grouping, an on-call engineer might receive dozens of notifications for what is essentially a single incident.
Alert grouping consolidates related alerts into a single incident, providing clarity on the nature of the outage while dramatically reducing notification noise.
For more information on configuring alert grouping and incident rules, see:
- OpsGenie: Automatically Create an Incident via Incident Rules - Configure rules to automatically create incidents from matching alerts, with built-in deduplication
- PagerDuty: Content-Based Alert Grouping - Group alerts based on matching field values like source, component, or severity
- PagerDuty: Time-Based Alert Grouping - Group all alerts on a service within a specified time window
Routing
Section titled “Routing”AxonOps provides a rich routing mechanism for notifications.
The current routing options are:
- Global - routes all notifications
- Metrics - alerts on metrics
- Backups - backups and restore events
- Service Checks - service checks and health checks
- Nodes - notifications raised from nodes
- Commands - notifications from generic tasks
- Repairs - notifications from Cassandra repairs
- Rolling Restart - notifications from the rolling restart feature
Each severity (info, warning, error) can be routed independently.

Errors per routing mechanism and severity levels
Section titled “Errors per routing mechanism and severity levels”Backup
Section titled “Backup”| Source | Severity | Description |
|---|---|---|
| Backup | Critical | Any error that is returned from the 3rd party remote location providers. |
| Backup | Warning | Clear local snapshots timed out |
| Backup | Warning | Unable to find local snapshot |
| Backup | Warning | Local backup process errors |
| Backup | Warning | Clear remote snapshot timed out |
| Backup | Warning | Remote backup process errors |
| Backup | Warning | Unable to find remote snapshot |
| Backup | Warning | Backup not triggered (Backups paused) |
| Backup | Warning | Failed to create backup |
| Backup | Warning | Failed to create remote config for backups |
| Backup | Warning | Create Cassandra snapshot failed |
| Backup | Warning | Snapshot request timed out |
| Backup | Warning | Cassandra node is inactive |
| Backup | Info | Local backup created successfully |
| Backup | Info | Backup deleted successfully |
Repair
Section titled “Repair”| Source | Severity | Description |
|---|---|---|
| Repair | Critical | Update repairs error, can be caused by tables being created or removed while a repair is running |
| Repair | Critical | Any error that is generated by Cassandra for a repair processes |
| Repair | Critical | Repair job is over 60% complete and the estimated time to completion is after gc_grace deadline |
| Repair | Warning | Repair job is over 40% complete and the estimated time to completion is after gc_grace deadline |
| Repair | Warning | Repair segment failed |
| Repair | Warning | Repair segment timed out |
| Repair | Warning | Cassandra repair error after n-amount of retries |
| Repair | Warning | Repair unit errors |
| Repair | Warning | Repair errors for nonexistent correlation ID |
| Repair | Warning | Repair request timed out after n-amount of attempts to connect to host |