An alert says a service is unhealthy. The person responding opens a dashboard, sees a gap, and has to decide whether the application failed or the evidence stopped arriving. A useful incident path should help answer both questions—and name the human who makes the next decision.

OpenTelemetry gives teams a common way to work with telemetry. This guide is Valen Systems' proposed exercise for turning that capability into a usable service runbook. It isn't a report from a Valen customer incident.

Graduation is background, not today's breaking news

OpenTelemetry's July 15, 2026 retrospective discusses the project's CNCF graduation in May 2026. It describes standardized telemetry specifications, APIs, language implementations, and the Collector, with ongoing work after graduation. It's a retrospective on a project milestone, not a September release announcement.

The milestone doesn't establish the behavior of a particular application's instrumentation or Collector configuration. Those remain things the operator needs to verify in the environment being reviewed.

Source: OpenTelemetry Has Graduated… Now what? — July 15 retrospective

Trace the configured path, not the component list

The Collector configuration documentation describes pipelines made of receivers, processors, and exporters. Defining a component isn't enough to enable it: it must be referenced in the service configuration. The order of processors in a pipeline matters. The documentation also describes a separate service telemetry setting for the Collector's own telemetry.

For the following exercise, use the configuration appropriate to your deployed version. Don't publish unredacted configuration, credentials, endpoint secrets, or customer records as evidence that the pipeline exists.

Source: OpenTelemetry Collector: Configuration

Choose one incident question

Start with a service and a question an operator would actually face. For example: did a submitted job finish, fail, or stop making progress? Agree on the observable outcome and the person responsible for deciding what happens next. Keep the first exercise narrower than a company-wide observability redesign.

Then identify the evidence you expect to use. Give each signal a job in the investigation. If nobody can explain which decision a proposed dashboard panel supports, leave it out of this first pass. You can add more once the path from symptom to decision works.

Walk a known event through the path

In an authorized test environment, choose a safe, non-sensitive example action with an identifiable time window. Record what the operator expects to see and where. Follow the evidence through the configured telemetry path to its destination. This is a proposed verification exercise, not a claim that every tool exposes the same intermediate detail.

Use synthetic identifiers rather than customer information wherever possible. If a step can't be observed directly, document how its behavior is inferred and what uncertainty remains. The exercise should produce a short evidence trail another operator can repeat, not a screenshot that only its author understands.

Ask what happens when the evidence is missing

Don't make the monitoring path's health an unstated assumption. Review how an operator would distinguish an application symptom from missing or delayed telemetry. Any deliberate interruption should be approved, bounded, and performed in a suitable test environment—not improvised against production. A tabletop discussion is still useful when a live exercise isn't safe, provided it's labeled accurately.

  • Who notices that expected evidence hasn't arrived, and what independent observation can they consult?
  • Which time window and service identity let the responder find the relevant records without searching unrelated customer data?
  • Which action is safe to take from the available evidence, and which action requires a human approval or additional check?
  • Who owns the application, who owns the telemetry path, and where does responsibility pass between them?
  • What stops a repeated response from continuing indefinitely, and who decides whether to roll back or escalate?
  • What evidence will show that the service is useful again, rather than merely that an alert has gone quiet?

Put the decision path into the runbook

Write the runbook around the incident question, the evidence locations, the bounded response, and the escalation owner. Keep observations separate from assumptions. Mark any untested step and the work needed to test it. A new operator should be able to tell where the documented procedure ends and judgment begins.

Valen Systems' Service Incident Runbook covers one service and up to three agreed scenarios, with response boundaries, escalation and rollback ownership, and a tabletop checklist. It isn't live incident response, monitoring installation, or an uptime promise. If an incident is active, use your existing incident responder first. Valen Systems' offer is follow-on documentation work, not an on-call response service.

Sources

  1. OpenTelemetry Has Graduated… Now what? — July 15 retrospectiveSource published July 15, 2026. Retrieved September 9, 2026.
  2. OpenTelemetry Collector: ConfigurationPublication date not stated. Retrieved September 9, 2026.

Put this to work.

Write a service incident runbook

Have a different situation? Talk with Valen Systems before choosing a scope.