Skip to main content
YavoraInnovations
Open menu

Practice 06

Observability Assessment

See your systems clearly without paying twice for telemetry you never query.

Engagement shapes

  • Cost and coverage review

    2 weeks

  • Full assessment

    4 to 6 weeks

  • Migration delivery

    Priced per scope

  • SLO workshop

    3 days

Who this is for

  • Platform and SRE leaders facing a renewal with a larger number on it
  • Teams carrying an alert volume nobody can act on
  • Organisations locked into one vendor with no portability plan

The problem

Telemetry volume grew every year. Signal did not. The bill tracks ingestion, and ingestion tracks whatever the last team instrumented by default.

Alert fatigue is the visible symptom. The underlying cause is usually that alerts were written against symptoms rather than against a service objective anyone agreed to.

By renewal time there is no portability story, so the negotiation happens with no leverage.

What we do

The workstreams in an engagement

01

Telemetry inventory

Coverage and gaps mapped by service and by signal, including what is collected and never queried.

02

Signal strategy

Logs, metrics, traces, and profiles assigned to purposes and retention tiers rather than collected uniformly.

03

OpenTelemetry migration path

Sequenced by service, with vendor portability assessed as an explicit outcome.

04

SLO design

Objectives and error-budget policy for the services that matter, written with the people who carry the pager.

05

Alert quality review

Noise, actionability, ownership, and routing, with a before-and-after volume estimate.

06

Incident and on-call review

Response process, escalation, and measured toil.

07

Cost and cardinality analysis

The specific drivers named, per service and per label, not a blended average.

08

Tooling evaluation

Consolidation options with the trade-offs stated, including the cost of staying put.

09

AI and LLM observability

Token cost, latency, quality signals, and trace linkage for model-backed services.

Deliverables

What lands on your desk

You keep all of it, including the method behind it, so the work can be repeated without us.

  • Coverage matrix by service and signal
  • SLO catalogue and error-budget policy
  • Alert rationalisation plan with a before-and-after volume estimate
  • OpenTelemetry migration plan with sequencing
  • Cost model with modelled savings and the assumptions behind them
  • Tooling recommendation with the trade-offs stated

Standards we work against

  • OpenTelemetry specification
  • Google SRE workbook SLO practice
  • Prometheus and OpenMetrics conventions

What we do not do

  • We do not recommend a platform migration where tuning the current one is cheaper.
  • We do not take commission from an observability vendor.

Engagement shapes

Three sizes, not one package

Durations are indicative and depend on estate size. Scope and price are fixed in writing before the engagement starts.

ShapeDurationWhat you get
Cost and coverage review2 weeksCost drivers, coverage gaps, and quick wins
Full assessment4 to 6 weeksCoverage matrix, SLO catalogue, alert plan, and cost model
Migration deliveryPriced per scopeOpenTelemetry pipeline in production
SLO workshop3 daysA first SLO catalogue written by your own teams

Related insights

Our thinking on this

Observability

SLOs before dashboards

Most observability spend is decided before anyone agrees what good looks like. That ordering is why the bill grows and the signal does not.

6 min read

Next step

Start with a briefing, not a proposal.

Thirty minutes on observability assessment. We will tell you whether we are the right firm for the problem, and who to talk to if we are not.