Begin with the service.
Identify what successful use of the service looks like. The most useful signals show whether users can complete the work the system exists to support.
Connect the supporting signals.
Resource utilisation, application logs and dependency checks help explain changes in service health. Keep their meaning documented and their context easy to find.
Give alerts a next step.
For each important alert, describe what should be checked first, who is responsible and how to tell that the problem is resolved.
Remove noise regularly.
Review alerts that were ignored or did not lead to a useful action. Adjust their conditions, improve the response instructions or remove them.
← All technical notes