Distributed systems trade the simplicity of a monolith for the complexity of coordinating dozens or hundreds of services. When a user's request fails after traversing an API gateway, two microservices, a message queue, and a database, finding the root cause without proper tooling is an exercise in frustration. OpenTelemetry provides a vendor-neutral, standardized approach to distributed tracing that gives you end-to-end visibility across your entire system.
This guide covers the practical architecture of implementing OpenTelemetry distributed tracing: SDK initialization, automatic and manual instrumentation, context propagation across service boundaries, sampling strategies for controlling costs, and backend integration for storing and querying traces. The patterns apply to any language, but the code examples focus on Node.js and TypeScript since they reflect common web service architectures. Whether you are tracing requests through a server-sent events pipeline or debugging latency in a zero-downtime migration, the fundamentals remain the same.
SDK Setup and Initialization
OpenTelemetry's Node.js SDK separates into three concerns: the API (interfaces for creating telemetry), the SDK (the implementation that processes telemetry), and exporters (backends that receive telemetry data). This separation means your application code depends only on the API, while the SDK and exporters are configured once at startup.
npm install @opentelemetry/sdk-node \
@opentelemetry/api \
@opentelemetry/sdk-trace-node \
@opentelemetry/exporter-trace-otlp-http \
@opentelemetry/resources \
@opentelemetry/semantic-conventions \
@opentelemetry/instrumentation-http \
@opentelemetry/instrumentation-express \
@opentelemetry/instrumentation-pg
The SDK must initialize before any other code runs. Create a dedicated instrumentation file that your application loads first, either through a --require flag or as an explicit import at the top of your entry point.
// instrumentation.ts
import { NodeSDK } from '@opentelemetry/sdk-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http';
import { Resource } from '@opentelemetry/resources';
import {
ATTR_SERVICE_NAME,
ATTR_SERVICE_VERSION,
ATTR_DEPLOYMENT_ENVIRONMENT_NAME,
} from '@opentelemetry/semantic-conventions';
import { HttpInstrumentation } from '@opentelemetry/instrumentation-http';
import { ExpressInstrumentation } from '@opentelemetry/instrumentation-express';
import { PgInstrumentation } from '@opentelemetry/instrumentation-pg';
import { BatchSpanProcessor } from '@opentelemetry/sdk-trace-node';
const traceExporter = new OTLPTraceExporter({
url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT || 'http://localhost:4318/v1/traces',
});
const sdk = new NodeSDK({
resource: new Resource({
[ATTR_SERVICE_NAME]: 'order-service',
[ATTR_SERVICE_VERSION]: process.env.APP_VERSION || '1.0.0',
[ATTR_DEPLOYMENT_ENVIRONMENT_NAME]: process.env.NODE_ENV || 'development',
}),
spanProcessors: [new BatchSpanProcessor(traceExporter)],
instrumentations: [
new HttpInstrumentation({
ignoreIncomingPaths: ['/health', '/ready', '/metrics'],
}),
new ExpressInstrumentation(),
new PgInstrumentation({
enhancedDatabaseReporting: true,
}),
],
});
sdk.start();
process.on('SIGTERM', () => {
sdk.shutdown()
.then(() => console.log('Tracing terminated'))
.catch((err) => console.error('Error shutting down tracing', err))
.finally(() => process.exit(0));
});
The Resource object attaches metadata to every span produced by this service. Service name, version, and environment are the minimum attributes needed to filter traces in any backend. The BatchSpanProcessor batches completed spans and exports them in bulk, reducing network overhead compared to exporting each span individually. The health and readiness endpoints are excluded from tracing because they generate high-volume, low-value spans that inflate storage costs.
To load this file before your application, add a --require flag to your start command.
node --require ./instrumentation.js dist/server.js
Automatic and Manual Instrumentation
Automatic instrumentation libraries monkey-patch popular frameworks and libraries to create spans without modifying your application code. The HTTP instrumentation captures every incoming and outgoing HTTP request. The Express instrumentation adds middleware spans showing route matching and handler execution. The PostgreSQL instrumentation wraps database queries with spans that include the SQL statement and execution time.
Automatic instrumentation covers the infrastructure layer, but business logic often needs manual spans to provide meaningful context. Use the OpenTelemetry API to create custom spans that describe domain-specific operations.
import { trace, SpanStatusCode, SpanKind } from '@opentelemetry/api';
const tracer = trace.getTracer('order-service', '1.0.0');
async function processOrder(orderId: string, items: OrderItem[]): Promise<Order> {
return tracer.startActiveSpan('process-order', {
kind: SpanKind.INTERNAL,
attributes: {
'order.id': orderId,
'order.item_count': items.length,
},
}, async (span) => {
try {
const inventory = await checkInventory(items);
span.addEvent('inventory-checked', {
'inventory.available': inventory.allAvailable,
'inventory.reserved_count': inventory.reservedCount,
});
if (!inventory.allAvailable) {
span.setStatus({
code: SpanStatusCode.ERROR,
message: 'Insufficient inventory',
});
throw new InsufficientInventoryError(inventory.unavailableItems);
}
const payment = await chargePayment(orderId, items);
span.addEvent('payment-processed', {
'payment.transaction_id': payment.transactionId,
});
const order = await createOrderRecord(orderId, items, payment);
span.setStatus({ code: SpanStatusCode.OK });
return order;
} catch (error) {
span.recordException(error as Error);
span.setStatus({
code: SpanStatusCode.ERROR,
message: (error as Error).message,
});
throw error;
} finally {
span.end();
}
});
}
The startActiveSpan method creates a new span and sets it as the active span in the current context. Any child spans created during execution, including those from automatic instrumentation, will be linked as children of this span. The addEvent calls record timestamped milestones within the span, which appear as annotations in trace visualizations. The recordException method captures error details following OpenTelemetry semantic conventions, including the stack trace.
For operations that cross async boundaries without automatic context propagation, you may need to explicitly pass context. The context API provides methods to capture and restore the active span context.
import { context, trace } from '@opentelemetry/api';
// Capture the current context
const currentContext = context.active();
// Later, in a callback or async handler
context.with(currentContext, () => {
const span = tracer.startSpan('deferred-work');
// This span is now a child of the captured context's span
span.end();
});
Context Propagation Across Services
Context propagation is the mechanism that links spans from different services into a single trace. Without it, each service produces isolated spans with no relationship to spans from other services. OpenTelemetry uses the W3C Trace Context standard by default, which defines two HTTP headers: traceparent and tracestate.
The traceparent header carries four fields: version, trace ID, parent span ID, and trace flags. A typical header looks like 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01. The trace ID is a 32-character hex string that uniquely identifies the entire trace. The parent span ID is a 16-character hex string identifying the calling span. The trace flags indicate sampling decisions.
// The HTTP instrumentation handles propagation automatically for fetch/http calls.
// For custom transports, propagate manually:
import { propagation, context, trace } from '@opentelemetry/api';
// Inject context into outgoing message headers
function sendMessage(queue: string, payload: unknown) {
const headers: Record<string, string> = {};
propagation.inject(context.active(), headers);
return messageQueue.publish(queue, {
body: payload,
headers, // traceparent and tracestate are now in headers
});
}
// Extract context from incoming message headers
function handleMessage(message: QueueMessage) {
const extractedContext = propagation.extract(
context.active(),
message.headers
);
return context.with(extractedContext, () => {
const span = tracer.startSpan('process-message', {
kind: SpanKind.CONSUMER,
attributes: {
'messaging.system': 'rabbitmq',
'messaging.destination': message.queue,
'messaging.message_id': message.id,
},
});
try {
// Business logic runs in the extracted context
processPayload(message.body);
span.setStatus({ code: SpanStatusCode.OK });
} catch (error) {
span.recordException(error as Error);
span.setStatus({ code: SpanStatusCode.ERROR });
throw error;
} finally {
span.end();
}
});
}
The key insight is that the HTTP instrumentation library handles propagation for standard HTTP calls automatically. You only need manual propagation for non-HTTP transports: message queues, gRPC (unless using the gRPC instrumentation), WebSockets, or custom protocols. For services that use Bun as a runtime, the OpenTelemetry HTTP instrumentation may require additional configuration since Bun's HTTP stack differs from Node's.
When integrating with systems outside your organization that may use different propagation formats, configure composite propagators to handle multiple header formats simultaneously.
import { CompositePropagator } from '@opentelemetry/core';
import { W3CTraceContextPropagator } from '@opentelemetry/core';
import { B3Propagator, B3InjectEncoding } from '@opentelemetry/propagator-b3';
const propagator = new CompositePropagator({
propagators: [
new W3CTraceContextPropagator(),
new B3Propagator({ injectEncoding: B3InjectEncoding.MULTI_HEADER }),
],
});
// Register globally
propagation.setGlobalPropagator(propagator);
Sampling Strategies
Sampling determines which traces are recorded and exported. In high-throughput systems processing thousands of requests per second, recording every trace is neither economically viable nor analytically useful. A well-designed sampling strategy reduces costs while preserving visibility into errors, slow requests, and representative traffic patterns.
OpenTelemetry provides several built-in samplers. The TraceIdRatioBased sampler uses a deterministic hash of the trace ID to select a configurable percentage of traces. Because the hash is deterministic, the same trace ID produces the same sampling decision in every service, ensuring that sampled traces are complete.
import { TraceIdRatioBasedSampler, ParentBasedSampler } from '@opentelemetry/sdk-trace-node';
// Sample 10% of traces, but respect parent decisions
const sampler = new ParentBasedSampler({
root: new TraceIdRatioBasedSampler(0.1),
});
const sdk = new NodeSDK({
// ...
sampler,
});
The ParentBasedSampler wraps the ratio-based sampler and ensures consistency across services. When a parent span is sampled, all child spans are also sampled. When a parent span is not sampled, children are not sampled. This prevents broken traces where some services sample a request and others do not.
For more sophisticated strategies, implement a custom sampler that considers request attributes. A common pattern is to always sample errors and slow requests while sampling a percentage of normal traffic.
import {
Sampler,
SamplingDecision,
SamplingResult,
} from '@opentelemetry/sdk-trace-node';
import { Context, SpanKind, Attributes, Link } from '@opentelemetry/api';
class AdaptiveSampler implements Sampler {
private baseSampler: TraceIdRatioBasedSampler;
constructor(private baseRatio: number) {
this.baseSampler = new TraceIdRatioBasedSampler(baseRatio);
}
shouldSample(
parentContext: Context,
traceId: string,
spanName: string,
spanKind: SpanKind,
attributes: Attributes,
links: Link[]
): SamplingResult {
// Always sample health-critical operations
if (attributes['http.route'] === '/api/payments') {
return { decision: SamplingDecision.RECORD_AND_SAMPLED };
}
// Always sample requests with debug flag
if (attributes['http.request.header.x-debug-trace'] === 'true') {
return { decision: SamplingDecision.RECORD_AND_SAMPLED };
}
// Drop noisy internal paths
if (spanName.startsWith('DNS') || spanName.startsWith('tcp')) {
return { decision: SamplingDecision.NOT_RECORD };
}
// Fall back to ratio-based sampling
return this.baseSampler.shouldSample(
parentContext, traceId, spanName, spanKind, attributes, links
);
}
toString(): string {
return `AdaptiveSampler(${this.baseRatio})`;
}
}
This approach ensures that business-critical operations like payment processing are always traced while routine health checks and DNS lookups are filtered out. The debug header provides an escape hatch for developers investigating specific requests in production.
The OpenTelemetry Collector
The OpenTelemetry Collector is a standalone binary that sits between your applications and your tracing backend. It receives telemetry from your services, processes it, and exports it to one or more backends. Running the Collector as an intermediary decouples your applications from specific backends and provides a centralized point for filtering, transforming, and routing telemetry data.
# otel-collector-config.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
timeout: 5s
send_batch_size: 1024
memory_limiter:
check_interval: 1s
limit_mib: 512
spike_limit_mib: 128
attributes:
actions:
- key: environment
action: upsert
value: production
- key: db.statement
action: hash
tail_sampling:
decision_wait: 10s
policies:
- name: errors
type: status_code
status_code: { status_codes: [ERROR] }
- name: slow-traces
type: latency
latency: { threshold_ms: 2000 }
- name: percentage
type: probabilistic
probabilistic: { sampling_percentage: 5 }
exporters:
otlphttp/tempo:
endpoint: http://tempo:4318
otlphttp/jaeger:
endpoint: http://jaeger:4318
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, tail_sampling, attributes, batch]
exporters: [otlphttp/tempo, otlphttp/jaeger]
The Collector configuration demonstrates several important patterns. The memory_limiter processor prevents the Collector from consuming unbounded memory under high load. The tail_sampling processor implements sampling decisions based on complete trace data, unlike head sampling which decides at trace creation time. Tail sampling always captures error traces and traces exceeding two seconds, while sampling five percent of normal traffic. The attributes processor hashes the db.statement attribute to prevent sensitive data like SQL parameters from reaching the backend.
Deploy the Collector as a sidecar alongside each service or as a gateway that receives telemetry from multiple services. The sidecar pattern provides better isolation and per-service buffering. The gateway pattern reduces the number of Collector instances and simplifies configuration management. Most production deployments use a two-tier architecture: sidecar collectors handle initial batching and forwarding, while a central gateway collector performs tail sampling and multi-backend routing.
Backend Integration and Querying
Traces need a backend for storage and querying. Popular open-source options include Jaeger, Grafana Tempo, and Zipkin. Managed services like Grafana Cloud, Datadog, and Honeycomb provide hosted backends with advanced query capabilities. The choice depends on your scale, budget, and existing monitoring infrastructure.
Grafana Tempo is a cost-effective choice for teams already using the Grafana stack. It uses object storage (S3, GCS, or Azure Blob Storage) for trace storage, making it significantly cheaper than alternatives that rely on Elasticsearch or Cassandra at scale.
# docker-compose.yml for local development
services:
otel-collector:
image: otel/opentelemetry-collector-contrib:0.96.0
volumes:
- ./otel-collector-config.yaml:/etc/otelcol/config.yaml
ports:
- "4317:4317" # gRPC
- "4318:4318" # HTTP
tempo:
image: grafana/tempo:2.4.0
command: [ "-config.file=/etc/tempo.yaml" ]
volumes:
- ./tempo.yaml:/etc/tempo.yaml
ports:
- "3200:3200"
grafana:
image: grafana/grafana:10.3.0
environment:
- GF_AUTH_ANONYMOUS_ENABLED=true
- GF_AUTH_ANONYMOUS_ORG_ROLE=Admin
ports:
- "3001:3000"
volumes:
- ./grafana-datasources.yaml:/etc/grafana/provisioning/datasources/datasources.yaml
With this setup, Grafana provides the visualization layer. TraceQL, Tempo's query language, lets you search traces by service name, operation, duration, and span attributes.
// TraceQL examples
// Find all traces from the order-service with errors
{ resource.service.name = "order-service" && status = error }
// Find traces where a database query took longer than 500ms
{ span.db.system = "postgresql" && duration > 500ms }
// Find traces where the payment processing failed
{ name = "process-payment" && status = error }
// Find traces that span multiple services
{ resource.service.name = "api-gateway" } >> { resource.service.name = "order-service" }
The >> operator in TraceQL finds traces where one span is an ancestor of another, which is invaluable for investigating cross-service issues. Combined with Grafana dashboards that surface RED metrics (Rate, Errors, Duration) derived from traces, this setup provides comprehensive observability without vendor lock-in.
Production Hardening and Best Practices
Moving from development to production requires attention to reliability, performance impact, and operational concerns. The tracing infrastructure must not degrade the performance of the systems it observes.
First, always use the BatchSpanProcessor in production. The SimpleSpanProcessor exports each span synchronously, adding latency to every operation. The batch processor buffers spans and exports them in bulk on a configurable schedule.
import { BatchSpanProcessor } from '@opentelemetry/sdk-trace-node';
const processor = new BatchSpanProcessor(exporter, {
maxQueueSize: 2048,
maxExportBatchSize: 512,
scheduledDelayMillis: 5000,
exportTimeoutMillis: 30000,
});
Second, set resource attributes that enable effective filtering. Beyond service name and version, include deployment environment, team ownership, and git commit SHA. These attributes turn traces from raw data into actionable intelligence.
const resource = new Resource({
[ATTR_SERVICE_NAME]: 'order-service',
[ATTR_SERVICE_VERSION]: process.env.APP_VERSION,
[ATTR_DEPLOYMENT_ENVIRONMENT_NAME]: process.env.DEPLOY_ENV,
'service.team': 'commerce',
'service.git.commit': process.env.GIT_SHA,
'service.instance.id': process.env.HOSTNAME,
});
Third, instrument your span creation with proper error handling. A tracing failure must never cascade into an application failure. Wrap all manual instrumentation in try-catch blocks, and configure the SDK's error handler to log rather than throw.
import { diag, DiagConsoleLogger, DiagLogLevel } from '@opentelemetry/api';
// Set diagnostic logger to warn level in production
diag.setLogger(new DiagConsoleLogger(), DiagLogLevel.WARN);
Fourth, be deliberate about what you include in span attributes. Avoid high-cardinality attributes like user IDs or request bodies in span attributes, as they inflate storage costs and can slow down backend queries. Use span events for supplementary data that you need occasionally rather than attributes that are indexed for search.
Finally, establish trace ID correlation with your logging system. Include the trace ID in every log line so that developers can jump between logs and traces when investigating issues. Most logging frameworks support custom fields that OpenTelemetry can populate automatically.
import { trace } from '@opentelemetry/api';
function getTraceContext() {
const span = trace.getActiveSpan();
if (!span) return {};
const spanContext = span.spanContext();
return {
traceId: spanContext.traceId,
spanId: spanContext.spanId,
traceFlags: spanContext.traceFlags,
};
}
// In your logger middleware
app.use((req, res, next) => {
const traceCtx = getTraceContext();
req.log = logger.child({
...traceCtx,
requestId: req.headers['x-request-id'],
});
next();
});
Distributed tracing with OpenTelemetry is not a weekend project. It requires thoughtful architecture from SDK initialization through backend selection. But the investment pays dividends every time you debug a cross-service issue in minutes instead of hours, or identify a performance bottleneck buried three services deep in your call chain. Start with automatic instrumentation and the Collector, add manual spans as you discover blind spots, and evolve your sampling strategy as your understanding of traffic patterns matures.