OpenTelemetry
What OpenTelemetry Is
Not a backend. Not a database. Not a dashboard. OpenTelemetry is a specification plus SDKs plus a wire protocol for producing telemetry and getting it out of a process. Where it goes afterwards is deliberately not its problem.
It is the merger of OpenTracing and OpenCensus, and it is now the second most active CNCF project after Kubernetes, which matters mainly as a statement about where the ecosystem has settled. Every serious vendor and open-source backend accepts OTLP.
The argument for using it is one thing: instrumentation is the expensive, slow, permanent part, and backends are the cheap, fast, temporary part. Instrumenting an application with a vendor SDK means the instrumentation has to be redone to change vendors, which in practice means never changing vendors. Instrumenting with OTel means the storage decision stays reversible, and the migration is a config change in a collector.
That is the entire pitch, and it is enough.
The Pieces
| Piece | What it is |
|---|---|
| Specification | Language-agnostic definition of the API, SDK behaviour, and data model |
| API | What application code calls. Stable, and a no-op unless an SDK is installed |
| SDK | The implementation: sampling, batching, processing, export |
| OTLP | The wire protocol: protobuf over gRPC or HTTP |
| Semantic conventions | Standard names for common attributes. The underrated part |
| Collector | A separate binary for receiving, processing and exporting. Its own page |
| Instrumentation libraries | Prebuilt spans for HTTP frameworks, DB drivers, queue clients |
The API/SDK split matters for libraries. A library can call the OTel API unconditionally; if the application has not configured an SDK, every call is a cheap no-op. This is what makes it safe for third-party libraries to be instrumented by default, which is how automatic instrumentation gets its coverage.
Signal Maturity
| Signal | Status |
|---|---|
| Traces | Stable. The mature part, and the original purpose |
| Metrics | Stable |
| Logs | Stable, but the newest. The SDK path is often a bridge from an existing logging library |
| Profiles | Development. The spec landed recently; expect churn |
Getting Spans Without Writing Code
Automatic instrumentation wraps known libraries at import or load time. For Python it is a wrapper command; for Java it is a JVM agent; for Node it is a require hook.
uv pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install # installs instrumentation for detected libs
OTEL_SERVICE_NAME=api \
OTEL_EXPORTER_OTLP_ENDPOINT=http://alloy:4317 \
OTEL_TRACES_SAMPLER=parentbased_traceidratio \
OTEL_TRACES_SAMPLER_ARG=0.1 \
opentelemetry-instrument python -m uvicorn app:mainThat produces spans for every HTTP request in and out, every database query, every cache call and every queue operation, with correct parent-child relationships and propagation, without touching the application.
Automatic instrumentation gets the structure right and knows nothing about the domain. It produces POST /api/orders and SELECT ..., never “this order was for a returning customer using a promo code”. The right approach is nearly always automatic instrumentation as the base plus a small number of hand-written spans and attributes where the business logic is.
Manual Instrumentation
from opentelemetry import trace
from opentelemetry.trace import Status, StatusCode
tracer = trace.get_tracer(__name__)
def process_order(order):
with tracer.start_as_current_span("process_order") as span:
span.set_attribute("order.id", order.id)
span.set_attribute("order.item_count", len(order.items))
span.set_attribute("order.total_cents", order.total_cents)
span.set_attribute("customer.tier", order.customer.tier)
try:
reserve_inventory(order) # child spans nest automatically
charge_payment(order)
except PaymentDeclined as e:
span.set_status(Status(StatusCode.ERROR, "payment declined"))
span.record_exception(e)
span.add_event("payment.declined", {"reason": e.code})
raise
span.set_status(Status(StatusCode.OK))
return orderstart_as_current_span puts the span in the context, so anything called inside it, including library code, creates children automatically. That implicit context is the whole mechanism, and it is also what breaks across a thread boundary or an unawaited task.
Span attributes are where high cardinality belongs. This is the inverse of the metric rule. An order ID, a user ID, a full URL, a query string: all fine as span attributes, because the cost model is bytes per span rather than a permanent new time series. Putting the interesting dimensions on spans is precisely what makes a trace able to answer the unanticipated question.
Events versus spans. An event is a timestamped annotation inside a span, with no duration of its own. Use a span for something that takes time and a child span for something that could be slow; use an event for a point-in-time fact, such as a cache miss or a retry decision.
Context Propagation
This is the mechanism that makes distributed tracing work, and the only part that genuinely breaks in practice.
Within a process, the current span lives in a context object, carried implicitly by a context variable. Across a process boundary, it must be serialised into the request.
traceparent: 00-7f3a1b2c3d4e5f60718293a4b5c6d7e8-00f067aa0ba902b7-01
tracestate: vendor1=opaque,vendor2=opaque
The W3C Trace Context headers are the default propagator, and B3 (Zipkin) and Jaeger formats are supported for compatibility. Instrumented HTTP clients inject and instrumented servers extract, automatically.
Where It Breaks
Message queues. Publishing to Kafka, RabbitMQ or SQS does not carry HTTP headers. The trace context must be injected into the message’s own metadata and extracted by the consumer. Instrumentation libraries do this for the common clients and nothing does it for a custom envelope format.
from opentelemetry.propagate import inject, extract
# Producer
headers = {}
inject(headers) # writes traceparent into the dict
queue.publish(body, headers=headers)
# Consumer
ctx = extract(message.headers)
with tracer.start_as_current_span("handle_message", context=ctx):
handle(message)Thread and task boundaries. A span started on one thread is not current on another. ThreadPoolExecutor and bare asyncio.create_task both lose it unless the context is captured and reattached, which the instrumentation does for common cases and not for hand-rolled concurrency.
Proxies and gateways that strip unknown headers. Rarer now, but a nginx or API gateway config that whitelists headers will drop traceparent and split every trace at that hop.
Batch and cron jobs. There is no incoming request to continue, so each run starts a new root trace. That is correct, and it means such jobs need a deliberate root span or they produce nothing at all.
The diagnostic is always the same: a trace that ends where it should continue, or two fragments that should be one. Work backwards from the boundary.
Semantic Conventions
The most undervalued part of the project. Semantic conventions are the agreed attribute names for common concepts.
http.request.method = "POST"
url.full = "https://api.example.com/orders?x=1"
http.response.status_code = 500
server.address = "api.example.com"
db.system = "postgresql"
db.query.text = "SELECT * FROM orders WHERE id = $1"
service.name = "api"
service.version = "1.4.2"
deployment.environment = "prod"
k8s.pod.name = "api-7d9f-x2k1"
Why it matters: every dashboard, alert and analysis tool that works out of the box depends on these names. Grafana’s service graph looks for service.name. A generic latency-by-route panel looks for http.route. Emitting endpoint instead of http.route means the ecosystem’s prebuilt everything ignores the data, and that cost is paid forever.
Use the conventions for anything they cover, and a clearly namespaced custom prefix for anything they do not: order.id, customer.tier, feature.flag.checkout_v2.
The conventions have changed, and versions matter. HTTP attributes were renamed during stabilisation (http.method became http.request.method, http.status_code became http.response.status_code). A fleet running mixed SDK versions emits both spellings, and dashboards need to handle it. The OTEL_SEMCONV_STABILITY_OPT_IN environment variable controls the transition in several SDKs.
Resource Attributes
Resource attributes describe the producer rather than the operation, and are attached to everything it emits.
OTEL_SERVICE_NAME=api
OTEL_RESOURCE_ATTRIBUTES="service.version=1.4.2,deployment.environment=prod,service.namespace=shop"service.name is the one that must always be set. Without it everything arrives as unknown_service, which is the most common misconfiguration in a new OTel deployment and makes the data nearly useless until fixed.
Configuration By Environment Variable
The SDKs are configured almost entirely through environment variables, which means configuration is a deployment concern rather than a code one.
| Variable | Purpose |
|---|---|
OTEL_SERVICE_NAME |
The service name. Always set this |
OTEL_EXPORTER_OTLP_ENDPOINT |
Where to send. Collector, or a backend directly |
OTEL_EXPORTER_OTLP_PROTOCOL |
grpc or http/protobuf |
OTEL_EXPORTER_OTLP_HEADERS |
Auth headers for a hosted backend |
OTEL_TRACES_SAMPLER |
always_on, always_off, parentbased_traceidratio |
OTEL_TRACES_SAMPLER_ARG |
The ratio, e.g. 0.1 |
OTEL_RESOURCE_ATTRIBUTES |
Comma-separated resource attributes |
OTEL_PROPAGATORS |
tracecontext,baggage by default; add b3 for compatibility |
OTEL_SDK_DISABLED |
true turns the whole thing off without removing it |
OTEL_METRIC_EXPORT_INTERVAL |
Metric push interval in ms, default 60000 |
OTEL_BSP_SCHEDULE_DELAY |
Span batch export delay in ms, default 5000 |
Port 4317 is OTLP over gRPC and 4318 is OTLP over HTTP, universally. The HTTP endpoint path is /v1/traces, /v1/metrics, /v1/logs, and OTEL_EXPORTER_OTLP_ENDPOINT should be the base URL without those suffixes.
OTel Metrics Versus Prometheus
Both work, and they differ in ways that surface at the worst time.
| Prometheus client | OTel metrics SDK | |
|---|---|---|
| Transport | Pull, over /metrics |
Push, OTLP |
| Instrument names | Counter, Gauge, Histogram, Summary | Counter, UpDownCounter, Histogram, Gauge, plus async observers |
| Temporality | Cumulative always | Cumulative or delta, configurable |
| Aggregation | In the query | In the SDK, configurable via views |
| Exemplars | Supported | Supported, automatic when a span is active |
Temporality is the trap. Prometheus assumes cumulative counters, which is what makes reset detection work. An OTel SDK configured for delta temporality sends the change since the last export, and feeding those into Prometheus produces nonsense: rate() on a delta series is meaningless. When exporting OTel metrics to Prometheus, set cumulative temporality, which is the default in most SDKs but not all, and is explicitly configurable via OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE.
The pragmatic position: OTel for traces everywhere, and the Prometheus client for metrics unless there is a specific reason to unify. Most stacks run both, and the Collector reconciles them.
Logs
The log signal is the newest, and the usual integration is a bridge rather than a new logging API. The application keeps using its existing logger, and a handler forwards records to the OTel pipeline with the current trace context attached.
import logging
from opentelemetry.sdk._logs import LoggerProvider, LoggingHandler
from opentelemetry.sdk._logs.export import BatchLogRecordProcessor
from opentelemetry.exporter.otlp.proto.grpc._log_exporter import OTLPLogExporter
provider = LoggerProvider()
provider.add_log_record_processor(BatchLogRecordProcessor(OTLPLogExporter()))
logging.getLogger().addHandler(LoggingHandler(logger_provider=provider))The point is trace correlation. A log record emitted inside a span carries that span’s trace ID automatically, which is what makes clicking from a trace to its logs work at all. Achieving the same thing with a file-based pipeline means manually injecting the trace ID into every log line, which works and is more fragile.
Practical Advice
Send to a Collector, not directly to a backend. Pointing SDKs straight at Tempo works and makes every future change a redeploy of every service. A Collector between them means the routing, sampling, filtering and backend choice are all config. This is the main argument of the Collector page.
Set service.name before anything else.
Start with automatic instrumentation, confirm traces arrive and connect across services, and only then add manual spans where the domain logic is.
Sample nothing at first. Sampling decisions made before seeing real volume are guesses, and a 10 percent sample during early adoption hides exactly the rare problems you are trying to find.
Check for broken propagation early, specifically across every queue and every non-HTTP boundary. It is much easier to fix while the system is small.
Where Next
- The OTel Collector for the pipeline that receives all of this.
- Alloy for Grafana’s distribution of that Collector.
- Tempo for where the traces land.