Direct Answer: What an OpenTelemetry Queue Failure Means
An OpenTelemetry Collector queue problem usually means that telemetry has reached a configured buffer, but the downstream exporter cannot send it fast enough. The queue may fill because the destination is slow, unavailable, rate-limited, or producing rejected data; alternatively, receivers and processors may be generating work faster than the exporter pipelines can process it. A full queue is therefore a symptom, not a single root cause, and increasing its capacity can postpone the failure without correcting the imbalance. The direct response should be to identify which pipeline owns the queue, inspect its utilization and rejection metrics, and then determine whether pressure originates at the receiver, processor, exporter, or destination. As of 29 September 2026, collectors may be self-hosted, operated by a cloud provider, or managed by an observability vendor, but the basic diagnostic model remains the same.
Also worth reading: How Do OpenTelemetry GenAI Conventions Work for LLM and Agent Observability in 2026? · How Do You Set Up OpenTelemetry Tracing for LLM Applications and AI Agents? · How Do Engineering Teams Implement OpenTelemetry for LLM Monitoring and Observability?
Do not treat every alert named “queue full” as an outage of the collector itself. If the process is healthy but the configured queue reaches a stated high-water mark, the collector is doing what the configuration instructed it to do: refusing new data rather than consuming unlimited memory. Conversely, if queue utilization is normal but telemetry disappears, the fault may be in sampling, filtering, a processor, authentication, or an exporter that silently discards records. The useful question is not merely whether the queue is full, but whether it was full of valid, unsent telemetry for long enough to violate the service’s reliability target. A sensible investigation target is to keep sustained utilization below roughly 80%, investigate at 70%, and treat 90–100% as a time-sensitive condition rather than a permanent normal state.
How the Collector Queue Actually Works
The Collector receives spans, metrics, and logs from one or more receivers and sends them through processors before exporters deliver them to a backend. Queueing commonly occurs inside an exporter, where records can wait for a batch to fill or for the next export attempt. A synchronous exporter may block instead of buffering, while a queued exporter separates temporary production from destination delivery and exposes metrics that make backlog visible. This separation improves resilience, but it also moves the risk from immediate send failure to delayed failure: accepted telemetry can remain in memory after the destination has already stopped responding correctly.
The commonly used memory-based queue has a configured queue size, and many distributed Collector distributions also expose a separate maximum queue size. The standard queueing exporter’s default queue size is commonly 1000 elements in current releases, but operators must verify the effective value because distributions, custom builds, and configuration overrides can change it. A batch may leave before it reaches the configured batch size because of a timeout, commonly 200 milliseconds in the queueing exporter defaults. Those numbers are defaults, not universal health thresholds, and they should be documented alongside any production override rather than assumed from a sample configuration.
A bounded queue protects the host by limiting memory consumption, while an unbounded queue would make backpressure and out-of-memory failure easier. The tradeoff is explicit: a larger queue absorbs a longer downstream outage but stores more unsent telemetry in RAM and can extend the period during which operators see stale data. Queueing also does not guarantee delivery. Data stored in a memory queue is lost if the collector terminates, and even persistent queueing is not a substitute for a durable event pipeline when the collector’s local disk is considered temporary. Understanding these mechanics prevents the common mistake of changing only queue_size when the real issue is an exporter retry loop that never makes progress.
A Practical Diagnostic Sequence
Start by defining the affected signal, pipeline, collector instance, and time window. A global graph labeled “OpenTelemetry” is not enough; identify whether the problem affects metrics, traces, logs, or all three, and determine which routing connectors lead to the affected exporter. Record the first bad timestamp, the deployment or configuration change immediately before it, the last known successful export, and the destination’s response. A precise window is especially important because an exporter that is only 5 minutes behind may require investigation, while a 2-hour delay can justify urgent action or a service-level escalation.
Next, inspect the collector’s own health endpoint and logs while preserving evidence from the failing instance. Look for receiver refusal, exporter send failures, authentication errors, timeouts, HTTP status codes, DNS failures, and messages about retry scheduling. Correlate them with queue size, queue capacity, and record/send counters exposed by the relevant exporter. If available, compare the same metrics across every replica rather than opening one pod randomly; a balanced 70% queue on five replicas may be stable, while an isolated 100% queue often points to affinity, configuration drift, or a destination shard problem.
Then separate three conditions: transient saturation, sustained saturation, and misleadingly healthy metrics. Transient saturation during a burst can be normal if the queue drains afterward and the destination’s ingestion quota is respected. Sustained saturation indicates that arrival throughput exceeds useful export throughput, and repeated full-capacity periods require capacity or routing changes. Misleading health occurs when telemetry is being rejected upstream, sampled away, routed to the wrong pipeline, or acknowledged by the destination in a way that does not match backend processing. At this stage, reproduce the problem with a controlled test signal and timestamped destination lookup; do not infer success merely because the collector logged a successful queue enqueue.
Reading Metrics, Logs, and Backpressure
A useful dashboard places queue behavior beside traffic and destination performance. Track incoming data rate, accepted and refused records, queue size, configured capacity, send attempts, successful sends, batch sizes, and retry or timeout logs. Calculate queue utilization as current size divided by capacity, then plot it over 15–30 minutes rather than relying on one sample. Also calculate the drain rate: if approximately 600 records per second are arriving and only 400 can be exported, the backlog grows by about 200 records per second even if every individual export attempt appears successful. This arithmetic often reveals the imbalance more clearly than a red queue gauge alone.
Watch for abrupt plateaus at the same value. A queue stuck at exactly 1000 with a configured size of 1000 may be full, but a queue frozen at 0 alongside rising refusal counters suggests a receiver-side problem. Queue growth can also be caused by oversized batches, expensive processors, memory-limit thrashing, a small exporter worker count, or an unnecessarily short backend timeout. Conversely, one very large batch can make memory usage unpredictable even if average throughput appears reasonable. Compare the failing period with at least two baselines: the last 7 days and the last known healthy release, because weekday traffic alone does not explain a post-deployment regression.
Metric names and labels vary by collector distribution, so search the actual telemetry rather than copying an assumed dashboard query. The upstream Collector is open source and documents its configuration and components, while managed services may rename metrics or expose only a subset. If a metric is absent, verify that the relevant telemetry pipeline is enabled and that the component is actually queued. Logs should provide timestamps, pipeline identity, error class, destination, and retry context without recording raw payloads, because payloads may contain secrets or regulated data. The best diagnosis combines quantitative trends with a single reproducible error, rather than choosing either dashboard graphs or log text in isolation.
Configuration Options and Their Tradeoffs
There is no universally correct queue setting. The right choice depends on the acceptable memory budget, expected burst size, maximum destination outage, and loss tolerance. A queue twice as large can absorb twice the unsent records under a simple workload model, but it also doubles the potential in-memory backlog; the actual effect can be larger if processors retain data elsewhere. Many teams begin with a bounded queue sized from measured traffic, reserve enough memory for concurrent batches, and define an operational threshold below 100%. Larger values should have an explicit rationale, such as surviving a 10-minute maintenance window, rather than being added merely to make an alert disappear.
| Feature | Memory Queue | Persistent File Queue | Direct or Synchronous Export |
|---|---|---|---|
| Storage location | Collector process memory | Collector host or volume | Direct connection to destination |
| Survives collector restart | No | Potentially, subject to recovery and disk settings | No queued recovery path |
| Main benefit | Low operational overhead and fast normal-path operation | Absorbs longer outages and selected restarts | Simple delivery path and immediate visibility into failures |
| Main risk | Memory growth and loss at restart | Disk pressure, latency, and operational complexity | Runtime backpressure that can affect collection |
| Best use | Short outages and high-throughput, memory-conscious services | Bounded outages where local buffering is acceptable | Low-volume data or controlled pipelines where blocking is acceptable |
Common Mistakes During Queue Troubleshooting
The most common mistake is increasing the queue before identifying the blocked destination. A larger queue can reduce immediate refusal while increasing memory consumption, process restarts, and the delay before data becomes visible to users. Another mistake is assuming that num_consumers, batch size, and queue size are independent performance controls. More workers can increase concurrent requests and provoke backend throttling; larger batches can improve compression but consume more memory and may exceed backend payload limits. Changes should therefore be made one at a time, with a defined measurement window and a rollback condition.
Teams also confuse “export succeeded” with “data is queryable.” An HTTP response can acknowledge ingestion before indexing is complete, and some integrations transform or drop records after acceptance. Conversely, a network timeout does not prove the backend failed; the request may have committed before the client lost the response. Use end-to-end verification with unique correlation values, backend ingestion records, and a realistic query delay. Finally, do not expose payloads while debugging, and do not test authentication by repeatedly sending production-scale bursts that trigger rate limits. Controlled synthetic data, redacted examples, and a single known-good record are safer and more informative.
When to Act and When to Wait
Act immediately when the queue remains at or near 100% for more than 5 minutes, the oldest queued item exceeds the team’s freshness objective, receiver refusal rises, or a collector pod is restarting from memory pressure. A 15-minute backlog may be acceptable for a low-priority internal telemetry stream during a planned maintenance window, but it may be unacceptable for security or billing data. Establish thresholds from service objectives instead of adopting a universal number. A useful starting policy is to investigate at 70% sustained utilization, escalate at 90%, and declare an incident when telemetry loss, memory pressure, or freshness loss threatens the agreed target.
Waiting briefly is reasonable when utilization is under 80%, drains within 2–5 minutes, no records are being refused, and a short deployment or autoscaling event is underway. Do not wait through a monotonically increasing queue, repeated identical authentication failures, DNS errors, or a growing receiver-refused metric. Those patterns indicate that retrying alone is unlikely to resolve the problem. During an active incident, prioritize containment: route nonessential signals to a less expensive or sampled path, shed low-priority telemetry, scale only the constrained tier, or temporarily disable an expensive processor if its output is not required. Preserve enough capacity for critical signals and record every mitigation so it can be reversed.
Cost, Capacity, and Vendor Variation
OpenTelemetry Collector software is open source, so the direct software license cost is generally $0; the real expense is compute, memory, storage, network egress, backend ingestion, and engineering time. A larger queue can reduce immediate loss but may require more memory or disk, while high-cardinality labels can make the downstream observability bill much larger than the queue itself. Managed collector and observability offerings may be priced per host, ingested volume, active series, spans, or log volume, so there is no defensible universal monthly price. Before buying capacity, measure bytes or records per second, peak retention in RAM or on disk, and destination cost per million events.
Vendor-managed OpenTelemetry services can reduce patching and operational work, but they may constrain queue depth, metric names, persistent storage, retry behavior, and export destinations. A self-hosted Collector offers configuration control and portability but transfers upgrade, security, capacity planning, and incident response to the operator. When a managed service offers an abstract queue status rather than detailed metrics, confirm whether the metric is queue utilization, export lag, ingestion delay, or simply an error count. Those are different signals. As of 29 September 2026, hybrid deployments are common, so the practical comparison should include behavior under an outage, data-loss guarantees, support response, and total monthly cost rather than a feature-count matrix alone.
A Durable Remediation Plan
A durable fix begins with a documented throughput model. Measure peak input rate, sustainable output rate, average encoded batch size, retry frequency, and the destination’s allowed request rate. Add enough safe capacity to handle the expected peak, usually with headroom above the observed maximum, but keep memory and cost constraints explicit. Then improve the constrained component: balance receiver and exporter pipelines, split high-volume traffic, use batching suited to the backend, remove unnecessary processors, reduce pathological cardinality, or route different signals to destinations with appropriate limits. If the backend itself cannot accept the traffic, increasing collector concurrency will only create more failed or throttled work.
Test the remediation under realistic failure conditions. At minimum, simulate a 5-minute backend interruption, a 30-minute interruption if the architecture claims longer resilience, a collector restart, and one replica loss. Measure accepted records, lost records, oldest data age, memory or disk high-water marks, and time to recover. A memory queue should be expected to lose unsent data on process loss, while a persistent queue should be tested for successful recovery rather than assumed to provide it. Finally, create an alert that reports sustained utilization and export lag, plus a runbook that names the destination, escalation owner, safe mitigation, and rollback procedure. This turns queue troubleshooting from a one-time change into a controlled observability practice suitable for AI-driven tutorials and production operations.