Observability programmes usually start with instrumentation, proceed to dashboards, and arrive at alerts. Service level objectives, if they appear at all, come last, written to describe whatever the dashboards already show.
Reversing that order changes what you collect, what you retain, and what you pay for. It is the single highest-leverage change available to most platform teams.
An objective decides what is worth collecting
Without an agreed objective, every signal is potentially useful, so teams collect everything at full fidelity and keep it for the longest retention the tool offers. That is a defensible decision in the absence of a stated goal, and it is exactly how ingestion bills grow.
An objective makes the question answerable. If the objective is a latency percentile on a checkout path, the telemetry that proves or disproves it is a small, specific set. Everything else can be sampled, aggregated, or retained briefly.
An objective makes alerts actionable
Symptom alerts fire on conditions that may or may not matter: elevated CPU, a queue depth, an error rate on one dependency. Each was added after an incident, and none carry a judgement about user impact.
Alerting on error-budget burn replaces a large number of these with a small number that always mean something. Volume typically falls sharply, and the alerts that remain are ones an on-call engineer can act on without first investigating whether the page was worth waking up for.
A practical measure of alert quality: what fraction of pages in the last quarter resulted in an action other than acknowledging them. Below half, the problem is the alert set, not the responder.
An objective gives cost work a limit
Cost reduction without objectives is a negotiation about risk with no shared reference. Every proposed reduction is met with the argument that the data might be needed during an incident.
With objectives, the argument resolves. Telemetry that proves an objective is protected. Telemetry that does not is a candidate for sampling or shorter retention, and the decision is a documented trade-off rather than a guess.
Writing the first ones
- 01Pick the three to five services whose failure is visible to customers. Ignore the rest for now.
- 02For each, write one availability objective and one latency objective, in the language of a user journey.
- 03Agree the target with the service owner, not with the platform team alone.
- 04Define the error budget and what happens when it is exhausted. An objective with no consequence is a dashboard label.
- 05Only then decide what telemetry is required to measure it, and at what fidelity.
Five objectives written this way do more for signal quality than another thousand dashboards, and they give the next renewal negotiation something to stand on.