Monitor
Monitor checks resources, certificates, cluster state, services, and synchronization, then optionally alerts or runs a selected service action on failure.
Where to find it: SDM root sidebar > Monitor.
Available to: root.
What this page does
The monitor turns operational expectations into repeatable checks. It polls the current state every 30 seconds in the page, while the configured check interval controls scheduled evaluation. Each item can be enabled independently and can alert only or request start, stop, or restart for an eligible service.
Before you start
Choose realistic thresholds, confirm email delivery, and understand which service an automatic action controls. Recovery actions can help with a stopped service; they cannot repair a full disk or a broken delegation plan.
Configuration
- Interval: 1 to 1,440 minutes; default 5.
- Email alerts: enable plus one or more comma-separated recipients; at least one address is required when enabled.
- Thresholds: CPU load, RAM percentage, disk percentage, and days before TLS expiry.
- Items: disk, CPU, RAM, TLS, cluster nodes, MariaDB, PowerDNS, SDM, and zone synchronization. The standalone deployment shows the local zone-sync row.
- Action on failure: none (alert only), start, stop, or restart where supported.
- Monitor service: root may start, stop, or restart the monitor itself.
Controls and fields
| Control, Field, Or Section | What It Does | What It Affects | Recommended Usage |
|---|---|---|---|
| Configure checks | Opens the Monitor configuration modal. | Allows changing alerting, enabled checks, failure actions, and thresholds. | Open it before production onboarding to define the monitoring baseline. |
| Enable email alerts | Switches monitor notifications on or off. | Controls whether configured recipients receive alert emails. | Enable only with valid recipients; disabled alerts still allow on-screen monitoring. |
| Alert destination email address(es) | One or more recipients for monitor alerts. | Determines where operational notifications are delivered. | Use a shared operations mailbox or ticketing address, not only a personal mailbox. |
| Services tab | Lists monitored service/system checks, current status, enabled state, failure action, and refresh control. | Controls which checks are active and what action SDM should take when a failure is detected. | Enable checks that match services the node actually runs. |
| Enabled checkbox | Turns a specific monitor item on or off. | Disabled checks no longer drive alerts/action behavior. | Disable only for intentionally unused services or during approved maintenance. |
| Action on failure | Selects the intended response such as no action, start, stop, restart, or status where supported. | Controls automated remediation behavior for service checks. | Use restart cautiously; avoid automated actions when diagnosing unstable services. |
| Refresh row | Runs or reloads one monitor check. | Updates that row status without changing configuration. | Use after restarting a service or renewing SSL to confirm the result. |
| Thresholds tab | Contains warning thresholds for CPU, RAM, disk, and SSL expiry. | Defines when resource/certificate states become warnings. | Tune thresholds to the customer workload so alerts are actionable. |
| CPU threshold | Load threshold for CPU warnings. | Triggers CPU pressure alerts when exceeded. | Set based on vCPU count and normal DNS query load. |
| RAM threshold | Memory usage percentage that becomes a warning. | Detects memory pressure before services fail. | Use a value that leaves enough headroom for updates and traffic spikes. |
| Disk threshold | Disk usage percentage that becomes a warning. | Protects DNS, logs, backups, and database storage from filling. | Keep conservative; full disks can break DNS updates and audit logging. |
| SSL threshold | Number of days before certificate expiry that becomes a warning. | Gives administrators time to renew or fix ACME problems. | Use enough lead time for customer change approvals. |
| Save changes | Stores monitor configuration. | Changes future monitoring/alert behavior. | Review enabled checks and recipients before saving. |
| Close | Closes the modal without applying unsaved edits. | Leaves current monitor configuration unchanged. | Use if you were only reviewing settings. |
How to read the result
| Status Or Message | Meaning | What To Check |
|---|---|---|
| Healthy | The monitored item is within expected state. | No action needed beyond normal monitoring. |
| Warning / degraded / partial | The check is working but one or more signals need attention. | Open details, refresh the row, then check related service/logs. |
| Down / failed / unavailable | The check could not confirm service health or detected a failure. | Inspect logs, service status, network reachability, and cluster peers. |
| Email alerts disabled | Alerts are not being sent even if monitor data is visible. | Enable alerts and set recipients if operational notifications are required. |
Practical guidance
- Use shared operations addresses for alert recipients.
- Keep automated restart actions conservative until normal service behavior is understood.
How to use it
- Open Configure checks.
- Set interval, recipients, thresholds, enabled items, and per-service action.
- Save, then confirm the monitor service is running.
- Select Refresh on representative rows and wait for the next page poll.
- Test alert delivery through an approved controlled condition rather than lowering a production threshold at random.
Result and next check
Enabled checks report current status, failures reach the intended recipients, and any configured service action is visible in state and logs. Threshold warnings should lead to the owning operational fix, not repeated restarts.
Theme color