Distributed Tracing: Following a Request Across Microservices
An engineer at a logistics startup got paged for a checkout flow that was intermittently taking eight seconds to respond, well past the threshold where customers started abandoning their carts. The request touched an API gateway, an authentication service, a pricing service, an inventory service, and a payments service, each with its own logs scattered across five different log groups.
Grepping through timestamps by hand to figure out which of the five services was really responsible for the delay took the better part of a night, and the answer, a slow downstream call inside the pricing service that only showed up under specific discount conditions, turned out to be something a single trace would have surfaced in minutes. That night is why distributed tracing exists.
Chasing a 500ms Mystery Across Twelve Services
Once an application splits into more than a handful of services, a single user-facing request can fan out into dozens of internal calls, and traditional logging stops being enough to explain what happened. Each service logs its own view of events, with no shared thread connecting them, so reconstructing the path of one request means manually correlating timestamps and guessing at causality, a process that gets exponentially harder as the number of services grows.
Distributed tracing solves this by attaching a single identifier to a request at the moment it enters the system and propagating that identifier through every subsequent call, so every service’s contribution to that one request can be reassembled afterward into a coherent timeline. Instead of five separate logs that might be related, you get one trace that shows exactly which service was slow, which calls happened in parallel, and where time really went.
Spans, Traces, and Context Propagation
A trace represents the full journey of one request through a distributed system. It’s composed of spans, where each span represents a single unit of work, an HTTP call, a database query, a function execution, with a start time, a duration, and metadata describing what happened.
Spans form a tree (or more precisely, in some models, a directed acyclic graph): a parent span for the overall API request contains child spans for each downstream call it makes, and those children can have their own children if they call further services. This structure is what lets a tracing UI render a waterfall diagram showing exactly how a request’s total latency breaks down across every service it touched.
- Trace ID: a unique identifier shared by every span belonging to one logical request.
- Span ID: a unique identifier for a single unit of work within the trace.
- Parent span ID: links a span to whichever span caused it, building the tree structure.
- Baggage: additional key-value context propagated alongside the trace, available to every span downstream.
Context propagation, passing the trace ID and current span ID from one service to the next, usually happens through HTTP headers following the W3C Trace Context standard (`traceparent`, `tracestate`), which most modern tracing tools and frameworks support out of the box.
For messaging-based communication, the same idea applies but the mechanism differs, since there’s no HTTP header to carry the context across a queue. Instead, the trace ID and span ID are embedded in the message payload or its metadata fields, and the consumer extracts them when it picks the message up, continuing the same trace even though the two sides of the exchange never directly connect to each other over the network.
This is one of the areas where teams most often lose track of a trace’s continuity, because it requires the message producer and consumer libraries to both cooperate on the convention, and a queue that simply passes opaque byte payloads through won’t do this automatically the way an HTTP client typically will.
Instrumenting Services With OpenTelemetry
OpenTelemetry has become the dominant standard for generating and exporting trace data, largely because it decouples instrumentation from any specific backend, you instrument your code once, and can send the resulting traces to Jaeger, Zipkin, a commercial APM tool, or several destinations at once, without changing application code.
A minimal manual instrumentation example in Python looks like this:
from opentelemetry import trace
tracer = trace.get_tracer(__name__)
def get_price(item_id):
with tracer.start_as_current_span("get_price") as span:
span.set_attribute("item.id", item_id)
price = pricing_client.fetch(item_id)
span.set_attribute("price.value", price)
return price
Most frameworks also support automatic instrumentation, which patches common libraries (HTTP clients, database drivers, web frameworks) to generate spans without any manual code changes, covering the majority of a typical request’s span tree with minimal setup effort. Manual instrumentation still earns its place for business-logic-specific spans, a complex pricing calculation, for instance, where automatic instrumentation has no way to know what’s worth capturing.
A practical rollout usually starts with the automatic instrumentation layer, since it delivers immediate value with almost no engineering time, and then layers in manual spans selectively as specific investigations reveal blind spots. A team debugging a slow checkout flow might discover that the automatic HTTP and database spans account for only sixty percent of the total request time, with the rest hidden inside an in-memory loop doing tax calculation across dozens of line items, exactly the kind of gap a manual span closes.
Over time, this incremental process converges on a trace tree detailed enough to answer most latency questions without engineers needing to add new instrumentation reactively during an active incident, which is a far worse time to discover a coverage gap than during a calm sprint planning session.
Sampling Strategies for High-Volume Systems
Capturing a full trace for every single request becomes prohibitively expensive at scale, both in terms of storage and in the overhead added to every request. Sampling decides which fraction of traces really gets recorded and stored.
- Head-based sampling: the decision to sample (or not) is made at the very start of a trace, before any spans exist, typically based on a fixed percentage or a probabilistic rule.
- Tail-based sampling: the decision is made after the full trace is complete, allowing rules like “always keep traces with errors” or “always keep traces slower than one second,” which head-based sampling can’t express since it doesn’t yet know the outcome.
- Adaptive sampling: the sample rate adjusts dynamically based on current traffic volume, keeping storage costs roughly constant even as traffic spikes.
- Priority sampling: certain request types (checkout, payment) get a higher sample rate than low-value traffic like health checks.
Tail-based sampling is generally more useful for debugging, since the traces that really matter, the slow ones, the failed ones, are exactly the ones it’s designed to keep, but it requires buffering complete traces before a keep-or-discard decision, which adds infrastructure complexity that head-based sampling avoids.
Jaeger and Zipkin in Practice
Jaeger, originally built at Uber and now a CNCF project, is one of the most widely deployed open-source tracing backends. It stores traces (commonly in Elasticsearch or Cassandra), offers a web UI for exploring individual traces and comparing latency across services, and integrates natively with OpenTelemetry as an export destination.
Zipkin, the older of the two major open-source options, pioneered much of the model that modern tracing tools still use, including the trace/span structure and header-based propagation. It remains in active use, especially in systems that adopted it before OpenTelemetry consolidated the ecosystem, though many new deployments now default to Jaeger or a commercial equivalent given the ecosystem’s broader current support.
- Storage backend choice: Elasticsearch offers flexible querying; Cassandra offers better write throughput at very high trace volume.
- Retention policy: traces are typically kept for days to weeks, not months, given their storage footprint at scale.
- Latency histogram overlays: both tools can show where a single trace falls relative to the overall distribution of recent request durations for the same endpoint.
- UI-driven investigation: engineers typically start from a slow or errored request and drill into its waterfall view to find the offending span.
- Service dependency maps: both tools can render a graph of which services call which, derived automatically from observed trace data.
Correlating Traces With Logs and Metrics
A trace tells you where time went; it rarely tells you the full story of why. The most productive tracing setups link each span to the specific log lines generated during that span’s execution, so an engineer can jump from “this span was slow” directly to the exact log output produced while it ran, without a separate manual search.
This correlation typically works by injecting the trace ID and span ID into structured log output as fields, so log aggregation tools (Elasticsearch, Loki, Datadog) can filter by trace ID and reconstruct the same timeline logs never could show on their own. Metrics complete the picture at a different resolution, while a trace shows one request’s journey, a metric like p99 latency per service shows the aggregate pattern across thousands of requests, and the two views are most useful together: metrics surface that something is wrong and roughly where, traces show exactly what happened in a specific instance of it going wrong.
The connective tissue between the two is often described as “exemplars”, a mechanism where a latency metric can point to a handful of specific trace IDs that represent the data points behind an unusual spike on a graph.
Instead of staring at a latency histogram and wondering which requests really drove the long tail, an engineer can click directly from the spike into one of the traces that produced it, collapsing what used to be a multi-step investigation into a single click. Not every metrics backend supports exemplars yet, but the pattern is becoming common enough in OpenTelemetry-compatible tooling that it’s worth checking for when evaluating a new observability stack.
Common Tracing Pitfalls

Instrumentation gaps are the most common problem teams run into, a context propagation header that isn’t forwarded through one particular internal call, silently breaking the trace into two disconnected pieces at that point, with no error or warning to indicate it happened. Asynchronous processing (a message queue, a background job) is a frequent culprit, since propagating trace context through a message payload requires deliberate effort that’s easy to skip.
- Broken context propagation: a service or library that doesn’t forward trace headers, fragmenting what should be one trace into several.
- Over-sampling in development, under-sampling in production: leading to a false sense of coverage that doesn’t hold up when it matters.
- Missing spans for expensive operations: automatic instrumentation covers common patterns, but a slow in-memory computation might not generate a span unless someone adds one manually.
- Cardinality explosion in span attributes: tagging spans with high-cardinality values like raw user IDs can blow up storage and indexing costs in some backends.
Clock synchronization is a subtler pitfall that only surfaces at scale. Spans generated on different hosts rely on each host’s local clock to timestamp their start and end, and even small amounts of clock drift between machines can make a waterfall diagram show a child span starting before its parent, or two spans that were really sequential appearing to overlap.
Most tracing backends tolerate this gracefully by rendering relative offsets rather than treating timestamps as perfectly authoritative, but engineers reading an unusual-looking waterfall should keep clock drift in mind as a possible explanation before concluding the causality itself is wrong. Running NTP or a similar time synchronization service consistently across the fleet reduces this problem by a wide margin, though it rarely disappears entirely in a large enough deployment.
Trace-Driven Debugging Workflows
The practical payoff of tracing shows up during incident response, where the difference between a five-minute diagnosis and a five-hour one often comes down to whether a trace exists for the affected requests. A mature workflow starts from an alert (elevated latency or error rate on a specific endpoint), pulls a sample of traces matching that condition, and looks for a pattern in the waterfall view, one particular span consistently taking longer than expected, or a particular downstream service showing up disproportionately in the slow traces.
Teams that build this habit into their on-call process tend to reach for traces before logs, reserving log search for the follow-up question of “why was this specific span slow” once tracing has already narrowed down which span that is. This ordering matters because logs without a starting point require guessing what to search for, while a trace hands you the exact service and time window to focus on.
Post-incident reviews benefit from the same artifacts. Attaching a representative trace to an incident writeup gives future readers a concrete, navigable record of what happened, rather than a paragraph of prose reconstructing it from memory and scattered log excerpts.
Some teams go further and build automated regression checks that compare a service’s trace shape over time, the number of downstream calls a given endpoint makes, or the typical span count in its critical path, flagging when a routine code change unexpectedly adds a new dependency or an extra round trip that wasn’t there in the previous release. This turns tracing from a purely reactive debugging tool into something closer to a continuous architectural health check, catching creeping complexity before it becomes the next 500ms mystery.
Final Thoughts
Distributed tracing earns its keep the first time it turns an overnight debugging session into a five-minute one, and after that, most teams wonder how they operated without it. The mechanics, trace IDs, spans, context propagation, are straightforward once you’ve seen them work, but the payoff depends entirely on consistent instrumentation across every service a request touches, since a single broken propagation link fragments the whole picture.
OpenTelemetry has made the tooling side of this far less painful than it used to be, decoupling how you instrument code from where the resulting data ends up. The investment is worth making before the eight-second checkout mystery happens, not after, because tracing is infrastructure you configure calmly in advance or scramble to bolt on during an incident, and only one of those is a good time.
Frequently Asked Questions
Does distributed tracing add noticeable overhead to requests?
Well-implemented tracing with reasonable sampling adds minimal overhead, typically a small percentage of request latency for span creation and export. The bigger cost is usually the storage and processing backend, which is why sampling strategy matters more than the instrumentation overhead itself.
Can tracing work across services written in different languages?
Yes. OpenTelemetry provides SDKs for most major languages, and because context propagation relies on standard HTTP headers rather than language-specific mechanisms, a trace can flow correctly from a Go service to a Python service to a Java service without special handling.
How long should trace data be retained?
Most teams retain full trace data for one to four weeks, driven by storage cost, and rely on aggregated metrics for longer-term trend analysis, since keeping every trace indefinitely rarely justifies its storage cost against how often old traces really get consulted.
What’s the difference between tracing and application performance monitoring?
Tracing is one component of APM. A full APM platform typically combines tracing with metrics, error tracking, and sometimes real user monitoring into a single tool, while tracing on its own focuses specifically on the request-level, cross-service latency breakdown.
Is tracing useful for a monolithic application?
Less so, since a monolith’s internal calls are function calls within one process rather than network calls across services, and a profiler is usually better suited to that kind of investigation. Tracing earns its value once requests truly cross service or process boundaries.
How do you get started with tracing on an existing system without a full rewrite?
Start with automatic instrumentation on the highest-traffic or most latency-sensitive service, export to a lightweight backend like Jaeger, and expand instrumentation outward to neighboring services incrementally, prioritizing the services most often implicated in past incidents.
