🔧 AI Nachrichten Debian is Voting on Whether to Allow AI-Assisted Contributions(23.08.2026 um 09:34 Uhr)
🔧 AI Nachrichten The Linux Kernel Is Approaching 2,000 CVEs Per Release(29.08.2026 um 20:00 Uhr)
⚠️ Malware / Trojaner / VirenCitrix Adds a Linux-Powered Escape Hatch For Compromised Windows PCs(30.08.2026 um 17:34 Uhr)
🔧 AI Nachrichten Debian is Voting on Whether to Allow AI-Assisted Contributions(23.08.2026 um 09:34 Uhr)
🔧 AI Nachrichten The Linux Kernel Is Approaching 2,000 CVEs Per Release(29.08.2026 um 20:00 Uhr)
⚠️ Malware / Trojaner / VirenCitrix Adds a Linux-Powered Escape Hatch For Compromised Windows PCs(30.08.2026 um 17:34 Uhr)

🔧 Programmierung 🕛 vor 3 Monaten 7 Min Lesezeit
0

Distributed Tracing 101: The Mental Model, the Standards, and Your First Pipeline

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

A request enters your system through an API gateway, hits an authentication service, queries a database, calls a payment provider, publishes an event to a message queue, and returns a response. When that request takes 4 seconds instead of 400 milliseconds, which service is responsible?



Without distributed tracing, you open five dashboards, compare timestamps in five different log streams, and try to reconstruct the request path from memory. With distributed tracing, you open one trace and see every hop, every duration, and every failure — in a single view.



Distributed tracing is the practice of propagating a unique identifier through every service that handles a request, recording the work each service does as spans, and assembling those spans into a trace that represents the request's complete journey.






The mental model: spans and traces



A span is a named, timed operation. "Query user table" is a span. "Call Stripe API" is a span. "Validate JWT" is a span. Each span records:




  • A name (what happened)

  • A start time and duration (how long it took)

  • A status (OK, error, or unset)


  • Attributes (key-value metadata: http.method=POST, db.statement=SELECT..., rpc.service=PaymentService)

  • A parent span ID (which span triggered this one)



A trace is a tree of spans rooted at the entry point. The root span represents the entire request. Child spans represent sub-operations. The parent-child relationships form a directed acyclic graph that mirrors the actual execution flow.




CODE
Trace: a]b2c3d4 (POST /api/v1/orders)
├── [12ms] Validate JWT
├── [340ms] Query order history
│ └── [320ms] PostgreSQL SELECT
├── [1,200ms] Call Stripe API
│ ├── [800ms] Create PaymentIntent
│ └── [380ms] Confirm PaymentIntent
└── [45ms] Publish OrderCreated event
└── [38ms] NATS publish






From this trace, you can immediately see that the Stripe API call dominates the latency (1,200ms out of ~1,600ms total). No log correlation, no dashboard cross-referencing, no guesswork.






Context propagation: the glue



Spans only form a trace if each service knows which trace it's participating in. This happens through context propagation — injecting the trace ID and parent span ID into the request headers, then extracting them on the receiving side.



The standard header format is
Collector + Query + UI
Elasticsearch, Cassandra, Kafka, Badger
CNCF graduated, battle-tested, flexible storage.
UI is functional but basic. No built-in metrics.


Zipkin
Monolithic or microservice
Cassandra, Elasticsearch, MySQL, in-memory
Simpler to deploy than Jaeger, smaller resource footprint.
Fewer features, smaller community, less active development.


Grafana Tempo
Distributed, object-storage-native
S3, GCS, Azure Blob
Cheapest at scale (no indexing). TraceQL is expressive.
Requires Grafana for visualization. Search depends on trace discovery (exemplars).


Datadog APM
SaaS
Managed
Zero operational burden. Unified with metrics and logs.
Expensive. Vendor lock-in.


Honeycomb
SaaS, columnar storage
Managed
Arbitrary-dimension queries. Excellent for high-cardinality.
Expensive at scale. Learning curve for BubbleUp queries.




For a detailed — they complement each other, they don't compete — see that guide.






Your first tracing pipeline



The fastest path to a working trace pipeline is: OTel SDK → OTel Collector → Jaeger. Here's a minimal setup.






1. Instrument your application



For a Node.js Express application:




CODE
npm install @opentelemetry/sdk-node @opentelemetry/auto-instrumentations-node \
@opentelemetry/exporter-trace-otlp-grpc









CODE
import { NodeSDK } from "@opentelemetry/sdk-node";
import { getNodeAutoInstrumentations } from "@opentelemetry/auto-instrumentations-node";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-grpc";

const sdk = new NodeSDK({
traceExporter: new OTLPTraceExporter({
url: "http://localhost:4317",
}),
instrumentations: [getNodeAutoInstrumentations()],
serviceName: "order-service",
});

sdk.start();






This auto-instruments HTTP, gRPC, database clients, and popular frameworks. Every incoming request creates a span. Every outgoing HTTP call creates a child span. Context propagation is automatic.






2. Run the OTel Collector



Use the config from our , Jaeger's Elasticsearch backend can run out of disk, and the network between your Collector and backend can partition. When any of these fail, traces are silently dropped — you don't notice until someone asks "why are there no traces for this incident?"



External monitoring closes the gap. A 30-second health check on your Collector's health endpoint and your Jaeger query service catches pipeline failures before the gap in your trace data becomes a blind spot. Set up these checks at .

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Bits und so #1021 (Passwort für Laufwerk)
1 Quelle
Bits und so #1022 (Wie Weißbier)
1 Quelle
KI-Agenten entdecken deutsches Wiki als Kommunikationskanal
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Distributed Tracing 101: The Mental Model, the Standards, and Your First Pipeline

Thematisch verwandte Begriffe: Distributed, Tracing, Mental, Model · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...