pkg/confighttp: the new keepalive section deprecates the flat keepalive fields in client and server config (#14020) — until v0.160.0, the client config exposed idle_conn_timeout, max_idle_conns, max_idle_conns_per_host, and disable_keep_alives as flat fields alongside the server config's idle_timeout and keep_alives_enabled. v0.160.0 deprecates all six flat fields in favor of a nested keepalive section; deprecated fields keep working exactly as before but emit a deprecation warning on construction, and setting a deprecated field together with the new keepalive section is an error. The migration path uses the new NewDefaultKeepaliveClientConfig and NewDefaultKeepaliveServerConfig helpers to build the nested section with default values. For an AI inference fleet whose Collector pipelines terminate OTLP from vllm, litellm, triton, otel-instrumented-agent, and a dozen self-hosted endpoints, this is the release where the keepalive config stops being a per-receiver copy-paste and moves to a single section per transport. The migration is non-breaking for now, but the deprecation timeline for these fields is the same one that aged out the pkg/exporterhelper and processor/batch histogram bucket boundaries in v0.157.0 — expect a Collector point release in the v0.163-0.165 window to remove the flat fields. Go 1.26 toolchain bump, drop 1.25 support (#15799) — the entire module set moves to go.mod Go 1.26.0, dropping support for 1.25.0. For teams that pin Collector CI to a Go matrix this is a CI toolchain bump, not a runtime change; for teams that ship a custom in-house Collector build, the build pipeline needs to follow. windows/amd64 is now tier 1; windows/arm64 is tier 2 (#15786) — Windows support tier formalization that mostly affects enterprise deployments where the Collector runs on Windows. pkg/exporterhelper: cache request size per sizer type so byte-sized batching no longer recomputes the whole batch on every consumed request (#12636) — A real perf fix. Until v0.160.0, with sizer: bytes the batcher's MinSize check called BytesSize() which ignored the cached size and re-walked the entire batch's protobuf on every consumed request, making accumulation O(n²) in the number of merged requests. v0.160.0 caches the size separately for the bytes and items dimensions and maintains it incrementally during merge/split, so the check is O(1). For a high-volume inference fleet whose OTLP exporter runs byte-sized batching against a ClickHouse or Tempo backend, this is the change that flattens the per-batch latency tail the observability 2026 pillar discusses as one of the operational scaling cliffs. processor/memory_limiter: only report component health status when the health state changes (#15751) — A health-reporting churn fix. Until v0.160.0, every internal state read on the memory_limiter produced a health-status log line, which on a steady-state Collector meant the health endpoint emitted the same OK line per scrape. v0.160.0 reports only on state changes. cmd/mdatagen: allow underscores in feature gate IDs (#15592) — A small but operationally meaningful fix for custom Collector components with underscored names (the convention the OTel contrib repo established in 2026). pkg/pprofile: add bounds checks to FromLocationIndices and switchDictionary (#15697) — A profile-merge crash fix that closes a real panic class. Until v0.160.0, FromLocationIndices did not bounds-check the location index, and a negative index decoded from an OTLP payload would panic the Collector via Profiles.MergeTo — a real remote-triggered crash class on a Collector ingesting profiles from an untrusted source. v0.160.0 returns an error instead. For an AI inference fleet that wires continuous profiling into the OTel pipeline (the same surface the probabilistic observability 2026 guide covers), this is the fix that closes the OTLP-profiles panic. pkg/service: route OpenTelemetry SDK-internal errors through the Collector's configured logger (#12378) — A log-shape fix. Until v0.160.0, an SDK-internal export failure (failed metric/log/trace export) was routed through the OTel SDK's default global error handler, which always printed to stderr via log.Print regardless of the configured log encoding. v0.160.0 routes those errors through the Collector's configured logger — meaning JSON-formatted Collectors now emit JSON for SDK errors, and a log-aggregation pipeline no longer sees mixed stdout/stderr streams. API: telemetry.Factory.CreateResource and telemetry.CreateResourceFunc now return (pcommon.Resource, string, error) to expose the resource's schema URL (#15129) — A breaking change on the experimental telemetry factory interface. The interface is experimental and cannot be implemented externally (it has an unexported method); callers using telemetry.WithCreateResource must update the function signature to return a third string value for the schema URL. For teams writing custom Collector distributions, this is a compile-time update. For everyone else, it is a non-breaking change because the interface is not externally implementable.
No migration is required for the user-facing config layer except the keepalive flat-fields deprecation, which is currently non-breaking (fields keep working with a warning) but should land before v0.163.0 to avoid hitting the removal window. The Go 1.26 toolchain bump means a CI matrix update on the same cycle. The byte-sized batching O(n²)->O(1) fix and the profile-merge crash fix are the two operationally invisible changes that move the needle most for AI inference fleets that ship the Collector in the hot path of an OTLP-to-Tempo or OTLP-to-ClickHouse pipeline. The observability 2026 pillar is the broader read on the AI/ML observability stack; the LLM monitoring stack tutorial walks through the OTel Collector + Prometheus + Grafana assembly this Collector cut feeds. For the AI semantic conventions side of the same OTel stack, the LLM context window optimization guide covers the token-side attributes the GenAI conventions emit alongside the trace data this Collector ingests.
otelcol_exporter_queue_batch_send_size_bytes and otelcol_processor_batch_batch_send_size_bytes have been replaced with a power-of-2 set spanning 128 B to 16 MiB (#15535, applied in both pkg/exporterhelper and processor/batch). Until v0.157.0, the old boundaries included many small sub-kilobyte buckets that were not useful for byte-scale payloads, and the exporter-side histogram topped out at 6000 bytes so nearly all real observations fell into the +Inf overflow bucket. The new boundaries are powers of two from 128 B to 16777216 (16 MiB), giving a meaningful distribution for real batch payload sizes (including small timeout-flushed batches) and keeping the two metrics directly comparable on the same dashboards. Dashboards or alerts that hard-code specific le values for these histograms will need to be updated. The big milestone is Phase 1 of the component configuration schema roadmap RFC (#14543) — every core component now ships a config.schema.yaml (debug/otlp/otlphttp exporters, otlp receiver, batch/memory_limiter processors, memory_limiter/zpages extensions) generated by the new schemagen tool and checked into the repo. That gives editors real autocompletion + validation on Collector configs without per-component bespoke schemas — the same shape of authoring-time win the JSON-schema-on-REST work in the Weaviate ecosystem produced for vector-database search responses. The other two high-impact changes are experimental resource detection for the Collector's own internal telemetry (#14311, enabled via service::telemetry::resource::detection/development::detectors) and the partial receiver-only reload feature gates — service.partialReload (Alpha) + service.partialReloadReceivers (Beta) — that together restart only receivers on config reload when non-receiver config sections are unchanged, avoiding unnecessary disruption to processors, exporters, and extensions (#5966). The bug fixes cluster around the gRPC client WaitForReady not being applied (#15615), the empty --feature-gates identifier panic (#15536), the cgroup-v2 slice-valued resource-attribute panic on startup (#15571), and the debug-exporter scope-index labelling bug that printed the parent resource's index instead of its own (#15541). For teams instrumenting LLM pipelines, v0.157.0 is the Collector release where the byte-scale histograms stop lying, the config schema ships with the binary, and reloads stop disrupting the processor + exporter chain.
OpenTelemetry CNCF graduation: what changed in July 2026
OpenTelemetry reached CNCF graduated status in May 2026, and the July 24 CNCF post from Adriana Villela of Dynatrace and Reese Lee of New Relic explains the operational result: OTel is production-ready, independently governed, and backed by more than 12,000 contributions from more than 2,800 companies. The project now gives platform teams a credible instrumentation boundary across traces, metrics, logs, and profiling.
Graduation does not freeze every gen_ai.* attribute. Individual semantic conventions still carry their own stability levels. Keep the convention version in deployment metadata, isolate experimental fields behind one helper, and test the emitted contract before Collector or SDK upgrades.
- Make OTLP the application boundary. Put backend-specific behavior in Collector exporters rather than vendor SDKs spread through inference services.
- Prove portability with a canary exporter. Send a small slice to a second backend and compare accepted spans, dropped attributes, query latency, and billable ingestion.
- Treat profiling as the fourth signal. Correlate CPU and allocation profiles with slow model-serving traces instead of operating a separate profiling island.
- Keep MCP telemetry adaptable. Record server, method, tool, duration, result, and retries, but hide experimental MCP attribute names behind an adapter.
A future OTel graduation spoke for AI/ML teams will include a minimal Collector profile, a backend migration drill, and a telemetry contract test; this in-place section is the live summary until the indexing gate clears. The main shift is commercial as much as technical: a verified OTel exit path lets a team test another backend without reinstrumenting every service.
OTel graduation → GA migration: what changes for AI semantic conventions in production
The 2026 graduation of OpenTelemetry from CNCF Incubating to Graduated is more than a milestone badge — it is the trigger for a GA migration that touches every AI semantic convention you ship today. The CNCF blog post "OpenTelemetry has graduated… now what?" (08-31) lays out the operational consequence: graduation locks in the project's independence, vendor support, and the long-term API stability guarantees platform teams need to invest in OTel as the production instrumentation boundary for AI/ML. The corpus's existing observability pillar covers the AI/ML observability stack in depth; this subsection covers the migration impact specific to the AI semantic conventions.
Three changes hit production AI telemetry pipelines in the graduation→GA window. First, semantic convention stability is now contractual. The gen_ai.* namespace and the in-progress gen_ai.agent.* extensions are now governed under the OTel SIG's stability policy — breaking changes to a stable attribute name require a deprecation cycle of at least one minor release plus a migration note in the spec repo. Dashboards or alerts that hard-code le bucket boundaries (the Collector v0.157.0 histogram change above is the canonical recent example) get the same treatment. Second, Collector configuration is now portable across vendors. The phase-1 config.schema.yaml rollout (#14543, in v0.157.0) means every core component ships an authoritative schema the editor can validate against — the same authoring-time win the Weaviate ecosystem produced for vector-database search responses. AI/ML teams that ran side-by-side exports to two backends during the Incubating phase can now make the second backend a primary without re-instrumenting — the schema-validated config is portable. Third, vendor SDKs converge on the OTel SDK as the instrumentation boundary. The graduation announcement explicitly names the SDK stability guarantee as the lever; vendor-specific SDKs that re-instrument inference are now the legacy path, not the recommended one.
What this changes for AI/ML teams running OTel today. The migration is not Big Bang — it is a four-step window the OTel SIG has been signaling since the May 2026 graduation vote. Step one (weeks 1-2): pin the OTel SDK + Collector versions in your deployment manifest, and add a CI check that the pinned version matches the OTel release notes' stability matrix. Step two (weeks 3-6): run a canary export to a second backend (the corpus's observability pillar covers the canary-export pattern in depth), compare accepted spans + dropped attributes + query latency, and decide whether the second backend is now your primary. Step three (weeks 7-10): contribute the AI-specific semantic conventions you have been carrying as vendor extensions (the gen_ai.agent.evidence.* namespace the evidence-packet analytics article uses is the canonical example) into the OTel SIG's stability review process. Step four (weeks 11-12): remove the vendor SDK layers your inference services shipped during the Incubating phase, and standardize on the OTel SDK as the application boundary.
What the migration does not change. OTel traces are still mutable. The gen_ai.* conventions still do not sign their spans. The Collector still does not give you a Rekor inclusion proof. For the audit surface — the surface the evidence-packet analytics article covers — the OTel graduation migration is a chance to clean up the operational telemetry layer, not the audit layer. The two layers compose: OTel traces are the live operational surface, evidence packets are the audit surface. The graduation migration sharpens the operational surface; the audit surface remains the schema-anchored, hash-chained, Rekor-anchored record the evidence-packet pattern produces. The TNS / CNCF August 2026 editorial framing is the canonical reference; the corpus's MCP + evidence-packet pillars compose with the OTel graduation migration into a complete AI/ML observability posture for late 2026.
Why Traditional Tracing Falls Short for LLMs
Standard distributed traces capture request-response pairs across service boundaries. LLM inference is different:
- Long, streaming responses — a single prompt generates hundreds of spans as tokens arrive sequentially
- Non-deterministic output — the same prompt can produce wildly different traces depending on sampling parameters
- Nested tool calls — an LLM agent can trigger cascading downstream API calls based on generated content
- Context propagation — prompt history, retrieved documents, and system instructions all affect output but aren't naturally captured in standard spans
OpenTelemetry's semantic conventions were extended in 2024-2025 to cover LLM-specific operations. Understanding these conventions is the first step to building meaningful observability for AI systems.
Setting Up the OTel SDK for LLM Instrumentation
The OpenTelemetry Python SDK provides first-class support for LLM instrumentation through the opentelemetry-instrumentation-openai and opentelemetry-instrumentation-litellm packages.
# Install dependencies
# pip install opentelemetry-api \
# opentelemetry-sdk \
# opentelemetry-exporter-otlp \
# opentelemetry-instrumentation-openai \
# opentelemetry-instrumentation-litellm
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.resources import Resource, SERVICE_NAME
# Initialize the tracer provider with service identity
resource = Resource.create({
SERVICE_NAME: "llm-inference-service",
"service.version": "1.0.0",
"deployment.environment": "production",
})
provider = TracerProvider(resource=resource)
processor = BatchSpanProcessor(OTLPSpanExporter(endpoint="http://otel-collector:4317"))
provider.add_span_processor(processor)
trace.set_tracer_provider(provider)
tracer = trace.get_tracer(__name__)
The resource attributes (service.name, deployment.environment) appear in every span and enable filtering in your observability backend.
Tracing OpenAI-Compatible API Calls
If your inference runs through an OpenAI-compatible endpoint (OpenAI, Azure OpenAI, local vLLM, Ollama, or any LiteLLM proxy), the openai instrumentation package captures spans automatically:
from openai import OpenAI
from opentelemetry.instrumentation.openai import OpenAIInstrumentor
# One-line instrumentation — wraps all OpenAI API calls
OpenAIInstrumentor().instrument()
client = OpenAI(api_key="sk-...")
with tracer.start_as_current_span("llm.summary-task") as span:
span.set_attribute("llm.prompt.template", "Summarize: {text}")
span.set_attribute("llm.prompt.variables.text", user_input[:200])
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "system", "content": "You are a technical summarizer."},
{"role": "user", "content": user_input}
],
temperature=0.3,
max_tokens=500,
)
# Extract response attributes for the span
span.set_attribute("llm.response.model", response.model)
span.set_attribute("llm.response.usage.prompt_tokens", response.usage.prompt_tokens)
span.set_attribute("llm.response.usage.completion_tokens", response.usage.completion_tokens)
span.set_attribute("llm.response.usage.total_tokens", response.usage.total_tokens)
span.set_attribute("llm.response.finish_reason", response.choices[0].finish_reason)
summary = response.choices[0].message.content
The openai instrumentation automatically creates child spans for each API call, capturing token counts, model selection, and latency. No manual span management required for the hot path.
Instrumenting Streaming Responses with Custom Spans
Streaming responses are the hardest case. The OpenAI SDK streams CompletionChunk events; each event contains a partial delta. Naively instrumenting each chunk creates an unmanageable span explosion. Instead, use a single parent span with chunk count attributes:
async def stream_completion(client, prompt: str, model: str = "gpt-4o"):
with tracer.start_as_current_span("llm.stream-completion") as span:
span.set_attribute("llm.prompt.length", len(prompt))
span.set_attribute("llm.model", model)
stream = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
stream=True,
stream_options={"include_usage": True}
)
full_response = ""
chunk_count = 0
# Consume the stream
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
content = chunk.choices[0].delta.content
full_response += content
chunk_count += 1
span.set_attribute("llm.response.chunk_count", chunk_count)
span.set_attribute("llm.response.full_length", len(full_response))
span.set_attribute("llm.response.first_token_latency_ms", first_token_ms)
return full_response
The stream_options={"include_usage": True} parameter (available in OpenAI SDK v1.35+) provides final token counts at the end of the stream, avoiding the need to accumulate counts during streaming.
Capturing LLM-as-Tool Calls with Semantic Conventions
Agentic AI systems use LLMs to decide when and how to call external tools. The OpenTelemetry semantic conventions for LLM operations (established in OTEP 0234) define standard attribute names for these interactions:
import json
def trace_tool_call(span, tool_name: str, arguments: dict, result: str, latency_ms: float):
"""
Record an LLM-initiated tool call using OTel semantic conventions.
https://opentelemetry.io/docs/specs/semconv/gen-ai/llm-semantic-conventions/
"""
span.set_attribute("gen_ai.system", "openai")
span.set_attribute("gen_ai.operation", "chat")
span.set_attribute("gen_ai.tool.name", tool_name)
span.set_attribute("gen_ai.tool.call.arguments", json.dumps(arguments))
span.set_attribute("gen_ai.tool.call.result", result[:500]) # Truncate for storage
span.set_attribute("gen_ai.tool.call.duration_ms", latency_ms)
The gen_ai.* prefix is the standard namespace for generative AI attributes in OpenTelemetry semantic conventions. These attributes enable you to filter and aggregate by tool name in Grafana, Honeycomb, or any OTel-compatible backend.
Context Propagation Across LLM Pipeline Stages
A production LLM pipeline typically involves multiple stages: prompt templating, retrieval augmentation, inference, and response post-processing. Each stage should propagate the same trace context to enable end-to-end latency analysis:
from opentelemetry.context import attach, detach
from opentelemetry.trace import SpanKind
def process_with_context(user_input: str, retrieved_docs: list[dict]):
# Extract the current trace context from the incoming request
# (assumes W3C trace context headers are present in the HTTP request)
context = extract_context_from_headers(request.headers)
token = attach(context)
try:
with tracer.start_as_current_span(
"llm.pipeline.full",
kind=SpanKind.INTERNAL,
) as span:
# Stage 1: Retrieval
with tracer.start_as_current_span("llm.pipeline.retrieval") as retrieval_span:
retrieval_span.set_attribute("retrieval.doc_count", len(retrieved_docs))
retrieval_span.set_attribute("retrieval.avg_doc_length", sum(len(d['content']) for d in retrieved_docs) // max(len(retrieved_docs), 1))
# Stage 2: Prompt assembly
with tracer.start_as_current_span("llm.pipeline.prompt-assembly") as prompt_span:
prompt = assemble_rag_prompt(user_input, retrieved_docs)
prompt_span.set_attribute("prompt.final_length", len(prompt))
# Stage 3: Inference (child of pipeline span, sibling of retrieval/prompt)
with tracer.start_as_current_span("llm.pipeline.inference") as inference_span:
response_text = call_llm(prompt)
inference_span.set_attribute("llm.response.length", len(response_text))
finally:
detach(token)
The W3C traceparent header propagates context across HTTP boundaries. If your pipeline involves async message queues (SQS, Pub/Sub), embed the serialized context in the message payload:
from opentelemetry.propagate import inject, extract
from opentelemetry.context import Context
def publish_to_queue(queue_url: str, payload: dict, parent_context: Context):
"""Inject trace context into a queue message for cross-service propagation."""
headers = {}
inject(headers, context=parent_context)
# headers now contains traceparent, tracestate for W3C propagation
payload["_trace_ctx"] = headers
sqs.send_message(QueueUrl=queue_url, MessageBody=json.dumps(payload))
How do you configure the OTel Collector for AI traces?
The OpenTelemetry Collector processes and exports your AI inference traces. A minimal production config for LLM traces:
# otel-collector-config.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
timeout: 1s
send_batch_size: 1000
# Filter out health-check spans
filter:
spans:
exclude:
match_type: strict
attributes:
- key: http.target
value: /healthz
# Redact PII from prompt content in traces
transform:
error_mode: ignore
trace_statements:
- context: span
statements:
- replace_pattern(attributes["llm.prompt.content"], "user_email=\".*\"", "user_email=\"[REDACTED]\"")
- replace_pattern(attributes["gen_ai.tool.call.arguments"], "\"api_key\":\".*\"", "\"api_key\":\"[REDACTED]\"")
exporters:
otlp/tempo:
endpoint: tempo:4317
tls:
insecure: true
prometheus:
endpoint: "0.0.0.0:8889"
namespace: llm_inference
const_labels:
service: llm-inference-service
service:
pipelines:
traces:
receivers: [otlp]
processors: [batch, filter, transform]
exporters: [otlp/tempo]
metrics:
receivers: [otlp]
processors: [batch]
exporters: [prometheus]
The transform processor is critical for compliance — prompts often contain user PII that shouldn't appear in your trace storage backend. If you are using the Collector's byte-scale histograms (the queue and batch send-size metrics the otelcol_exporter_queue_batch_send_size_bytes and otelcol_processor_batch_batch_send_size_bytes histograms surface), note that v0.157.0 replaced the old boundaries with a power-of-2 set spanning 128 B to 16 MiB, so any dashboard or alert that hard-codes the old le values will need to be updated alongside this config. For the broader observability stack around the Collector — pulling in the right sampling, tail-based decisions, and SLO surfaces — the observability 2026 guide walks through the trace + metric + log mix that complements the Collector, and the Prometheus vs Grafana 2026 comparison covers the metric-pipeline side. If you are running the Collector on Kubernetes and want to control the cost of the metric pipeline at scale, the resource-quota and cgroup-v2 GOMAXPROCS interactions that v0.157.0's resource-detector fixes interact with are the same family of behavior the Kubernetes GPU operators guide covers for compute-bound workloads.
Key Metrics to Capture from LLM Pipelines
Beyond traces, your LLM pipeline should emit these metrics for capacity planning and cost monitoring:
| Metric | Type | Description |
|---|---|---|
llm.request.duration | Histogram | End-to-end latency from prompt receipt to final token |
llm.tokens.prompt | Counter | Total prompt tokens processed |
llm.tokens.completion | Counter | Total completion tokens generated |
llm.tokens.cost | Counter | Estimated API cost in USD |
llm.stream.first_token_latency | Histogram | Time to first token (TTFT) |
llm.error.rate | Gauge | Rate of API errors / rate limit hits |
llm.tool.call.duration | Histogram | Per-tool call latency |
Emit these with the OpenTelemetry Metrics API:
from opentelemetry import metrics
meter = metrics.get_meter(__name__)
token_counter = meter.create_counter(
name="llm.tokens.total",
description="Total tokens processed",
unit="1",
)
# Record token usage after each response
token_counter.add(
response.usage.prompt_tokens,
{"token.type": "prompt", "model": response.model}
)
token_counter.add(
response.usage.completion_tokens,
{"token.type": "completion", "model": response.model}
)
Grafana Dashboard for LLM Pipeline Observability
With traces flowing into Grafana Tempo and metrics into Prometheus/Grafana Cloud, here's the minimum viable dashboard layout for LLM pipeline observability:
Panel 1: Request Volume + Error Rate
sum(rate(llm_inference_requests_total[5m])) by (model)
| vs |
sum(rate(llm_inference_errors_total[5m])) by (model)
Alert threshold: error rate > 1% triggers PagerDuty.
Panel 2: Token Cost by Model
sum(increase(llm_tokens_total[24h])) by (model, token.type) * token_price_per_1k
This panel is essential for FinOps — you should know within 5% accuracy what each model costs per day.
Panel 3: P50/P95/P99 Latency Heatmap
histogram_quantile(0.95, sum(rate(llm_request_duration_bucket[5m])) by (le, model))
P99 > 30s for streaming models is a signal to investigate cold start issues or rate limiting.
Panel 4: Tool Call Frequency
sum(rate(gen_ai_tool_calls_total[5m])) by (tool.name)
Spikes in tool call frequency indicate your agent is entering loops or encountering ambiguous prompts.
Panel 5: First Token Latency Distribution
histogram_quantile(0.50, sum(rate(llm_first_token_latency_bucket[5m])) by (le, model))
TTFT > 5s on a streaming endpoint will feel unresponsive to users. Establish baselines per model version.
Grafana Cloud is the easiest way to get started with LLM observability — native OTel support, pre-built dashboards for LLM traces, and pay-as-you-go pricing for inference-heavy workloads.
Common Pitfalls in LLM OTel Instrumentation
Pitfall 1: Capturing full prompt content in every span. Prompts can be thousands of tokens. Storing the full prompt in every span quickly inflates your trace storage costs. Instead, record llm.prompt.length and llm.prompt.template, storing the actual prompt in a separate document store keyed by a hash.
Pitfall 2: Missing sampling for high-volume endpoints. If your inference service handles 10,000 requests/minute, 100% trace sampling is expensive. Use Tail-Based Sampling (available in OTel Collector) to capture 100% of error traces but only 1-5% of successful traces.
Pitfall 3: Forgetting to propagate context through async boundaries. LLM pipelines often use background queues (Celery, SQS, RabbitMQ). If you don't serialize and propagate the W3C trace context, your traces will show up as separate unconnected spans rather than a single pipeline trace.
Pitfall 4: Ignoring model version drift. The same model name can return different versions over time (GPT-4-Turbo's underlying weights are updated). Add model.version or model.revision to your resource attributes if your inference provider exposes this information.
Putting It Together: A Complete Instrumented Pipeline
Here's the full pattern — a FastAPI endpoint with complete OTel instrumentation including context propagation, streaming support, and cost attribution:
from fastapi import FastAPI, Request
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
from opentelemetry.propagate import extract
from opentelemetry import trace, metrics
app = FastAPI(__name__)
FastAPIInstrumentor.instrument_app(app)
meter = metrics.get_meter(__name__)
request_counter = meter.create_counter("http.llm.requests", "LLM HTTP requests")
token_histogram = meter.create_histogram("llm.tokens", "LLM token usage")
@app.post("/v1/chat")
async def chat(request: Request, body: ChatRequest):
ctx = extract(dict(request.headers))
with trace.get_tracer(__name__).start_active_span(
"chat",
context=ctx,
attributes={"http.method": "POST", "llm.model": body.model}
) as span:
# Call LLM (instrumented automatically via OpenAIInstrumentor)
response = client.chat.completions.create(
model=body.model,
messages=body.messages,
stream=body.stream,
)
if body.stream:
return StreamingResponse(stream_tokens(response, span))
else:
return format_response(response)
The combination of automatic instrumentation (OpenAIInstrumentor handles the API calls) and manual spans (you control business logic boundaries) gives you complete visibility without excessive overhead.
Conclusion
OpenTelemetry has matured into the standard for LLM observability, but applying it to AI inference pipelines requires understanding its AI-specific semantic conventions. The key patterns — streaming span aggregation, gen_ai.* attributes for tool calls, W3C context propagation across async boundaries, and PII redaction — are what separate production-grade AI observability from toy implementations.
Building this instrumentation now pays compound dividends as your inference volume grows: the same traces that help you debug a single incident become the dataset for capacity planning, cost attribution, and model selection decisions.
Start with the OpenAI instrumentation package (one line of code), add custom spans for your pipeline-specific logic, and wire up a Grafana dashboard for the metrics that matter to your team. From there, expand into tail-based sampling and cost attribution as your observability needs mature.