# How Do You Configure OpenTelemetry Collector Queues Without Losing Telemetry?

aitutorialmaker.com · October 1, 2026

> What OpenTelemetry Queue Configuration Actually Controls OpenTelemetry Collector queue configuration determines where telemetry waits between...

## What OpenTelemetry Queue Configuration Actually Controls

OpenTelemetry Collector queue configuration determines where telemetry waits between collection and export. It is most often configured through the sending_queue block in an exporter, while the Collector’s memory_limiter processor provides a separate, system-wide safeguard against excessive memory use. These queues are temporary buffers, not durable message brokers: data held in memory can disappear during a crash, restart, forced shutdown, or prolonged backend outage. As of 1 October 2026, the exact components available depend on the Collector distribution, so contrib, Kubernetes, and vendor-specific builds do not necessarily expose identical processors. The controlling settings include enabled, queue_size, num_consumers, and retry_on_failure; sensible defaults may be adequate for a development environment, but production traffic should be measured before those values are treated as universally safe. Queue configuration therefore answers a narrow operational question: how long and how much telemetry the Collector may hold before it must block, retry, reject, or drop data.

**Also worth reading:** [How Can You Use AI to Create Simple Tutorials Without Losing Accuracy?](https://aitutorialmaker.com/knowledge/how_can_you_use_ai_to_create_simple_tutorials_without_losing_accuracy.php) · [How Do You Build an AI Tutorial Video Pipeline Without Losing Quality?](https://aitutorialmaker.com/knowledge/how_do_you_build_an_ai_tutorial_video_pipeline_without_losing_quality.php) · [How Do You Monitor OpenTelemetry Queue Metrics and Diagnose Backlogs in 2026?](https://aitutorialmaker.com/knowledge/how_do_you_monitor_opentelemetry_queue_metrics_and_diagnose_backlogs_in_2026.php)

The central rule is to make loss behavior explicit. Enabling a larger queue increases burst tolerance but also increases the amount of telemetry that may be at risk when a process terminates. A queue of 2,048 items is the common OpenTelemetry default in many exporter implementations, but 2,048 records at 500 bytes each is only about 1 MiB; the real memory requirement can be much larger after batches, attributes, compression, and internal overhead are considered. Likewise, 10 concurrent consumers is a common starting point, not a guarantee of ten times the throughput, because each consumer is constrained by the exporter protocol and destination. The correct settings are derived from observed item size, export latency, acceptable loss tolerance, and available memory—not from a copied example alone.

## Recommended Collector Settings for Most Production Systems

A robust starting point is to enable the exporter queue, set a documented queue size, and inspect how the Collector behaves during an outage before increasing concurrency. A typical YAML block is sending_queue: {enabled: true, queue_size: 5000} on an OTLP exporter, together with retry_on_failure: {enabled: true, initial_interval: 5s, max_interval: 30s, max_elapsed_time: 300s}. The queue size should be treated as an engineering budget: 5,000 queued items might consume tens or hundreds of megabytes if records are large. The retry window should match the recovery objective; a five-minute limit can prevent the Collector from retaining data indefinitely, but it also makes a longer backend incident more likely to culminate in drops. Always confirm the field names and supported options against the documentation for the exact Collector release and distribution in use.

Concurrency is a second decision. Starting with 2 to 10 consumers can prevent a single slow request from serializing all exports, but adding consumers too quickly may overload a backend, increase connection counts, or consume additional CPU and memory. Benchmark at least 2, 4, and 10 consumers under realistic batch sizes rather than assuming that the largest value performs best. If the Collector must export to a rate-limited SaaS service, local queueing can protect the process briefly, but the destination’s documented ingest and burst limits remain the final constraint. Metrics such as otelcol_exporter_queue_size, otelcol_exporter_send_failed_metric_points, and refused-send counts should be watched, although metric names and stability can differ by release. Queue tuning is effective only when paired with alerts and traces of failed requests.

## Memory Limiting Versus Export Queue Buffering

A sending queue absorbs bursts in the data path, whereas the memory_limiter processor protects the Collector by applying backpressure or refusing telemetry when memory crosses configured thresholds. The memory_limiter commonly uses a soft limit around 80% of the system’s allocated container memory and a hard limit around 90%, leaving headroom for telemetry itself, the Go runtime, and non-Collector processes. This is not a replacement for exporter queues: a memory limiter can push pressure upstream, while an exporter queue can continue retaining data. Combining both features is usually sensible because they enforce different boundaries. If the hard limit is set too close to the container limit, the Linux kernel’s OOM killer may terminate the process before the processor can react cleanly.

| Feature | Exporter sending queue | Memory limiter processor | External message broker |
| --- | --- | --- | --- |
| Primary purpose | Buffer records before export | Control Collector memory pressure | Provide durable, decoupled delivery |
| Default location | Collector process memory | Collector-wide control point | Separate process or service |
| Survives Collector restart | Usually no | No | Potentially yes |
| Typical implementation | queue_size and num_consumers | Soft and hard memory limits | Kafka, RabbitMQ, or cloud queue |
| Best fit | Short bursts and transient latency | Protecting constrained hosts | Strict separation and longer outage tolerance |
| Main cost | Heap memory and possible data loss | Upstream pressure or dropped telemetry | Operations, latency, and infrastructure expense |

The distinction matters because queues do not make delivery end-to-end reliable by themselves. An OpenTelemetry Collector is a telemetry gateway, not a general-purpose event ledger, and most standard OTLP exporter queues are in-memory structures. A durable broker can sit between collection and export when loss tolerance is strict, multiple consumers need independent scaling, or very long outages are expected. That approach adds serialization, network, storage, and operational complexity, so it is usually excessive for low-value development logs. It becomes defensible for regulated audit events, billing-related telemetry, or production traces where the business has agreed that losing a five-minute period is unacceptable.

## A Practical Step-by-Step Configuration Method

First, measure the uninstrumented baseline by recording ingestion rate, average item size, export latency, CPU, heap use, and restart frequency for at least one representative traffic cycle. For example, a service generating 20,000 spans per second creates far different queue pressure from one producing 200 spans per second, even if the exporter names are identical. Next, establish a data-loss budget with the service owner, such as accepting less than 0.1% of spans during a short incident while never allowing Collector memory above 80% of its allocation. Those are policy targets rather than OpenTelemetry guarantees, and they should be translated into alert thresholds. Only then choose a queue size using the measured item size and tolerated outage duration rather than rounding by intuition.

After enabling queueing, test normal traffic, a destination slowdown, a complete network outage, and a Collector restart. A network test that only blocks new connections may not reproduce the behavior seen when an endpoint accepts requests but responds slowly. Increase num_consumers in measured increments, such as 2, 4, 8, and 10, and compare successful exports per second, p95 latency, memory, and failure count. If larger consumer counts do not improve throughput, retain the lower value because they may simply create more contention. Finally, document a shutdown policy: some deployments deliberately use a long Kubernetes termination grace period to drain the queue, but this is useful only if the process receives a graceful termination signal and the queue contains records that can be flushed. A SIGKILL or node failure defeats the drain period.

A useful production test should include telemetry-cardinality extremes. Ten small log records may consume less memory than two records containing long error messages, exception stack traces, or hundreds of resource attributes. Queue capacity measured only in items is therefore incomplete; monitor bytes and heap allocation where the runtime and collector build permit it. Also test what happens when an upstream SDK has a full sending queue. The SDK and Collector are separate buffering systems, so increasing one does not automatically protect the other. Preserving important telemetry requires a chain of tested policies across application SDKs, Collectors, gateways, and the backend.

## Common Queue Configuration Mistakes

The most frequent mistake is copying a large queue_size without calculating memory. Doubling a queue from 2,048 to 4,096 items does not merely double a small fixed buffer; it can materially raise heap use, garbage-collection work, and restart exposure. Another common error is enabling retries forever with no max_elapsed_time, allowing stale telemetry to remain while newer operational information cannot move efficiently. Turning on both exporter retry and processor retry without understanding the interaction may also multiply delays, so their timeouts should be reviewed together. Queue settings are not a substitute for a memory limiter, and a memory limiter is not evidence that a backend outage is harmless.

Batch size is frequently confused with queue size. A queue can hold thousands of pending items, while the exporter sends them according to its batching behavior and protocol limits; increasing one does not necessarily change batch payloads in the way an operator expects. Setting consumer concurrency far above available network throughput can generate retries and duplicate pressure without improving successful delivery. It is also a mistake to assume the deprecated or removed batching behavior of one Collector release applies to another. Distribution-specific changes are important because the OpenTelemetry Collector core contains fewer components than the opentelemetry-collector-contrib distribution, and even contrib components evolve over time.

Finally, teams often monitor process health but not export correctness. A healthy Collector can be up while silently dropping telemetry, and a growing queue can be mistaken for healthy buffering even when the oldest records will expire before export. Record dropped-item counts, refused items, queue utilization, export duration, and backend ingestion independently. Change configuration through version control, validate it in a staging deployment, and retain the previous manifest for rapid rollback. These controls matter because queue behavior at a backend outage boundary is exactly where production-only assumptions tend to fail.

## When to Increase Capacity, Change Architecture, or Accept Loss

Increase the queue when short bursts are frequent, the destination normally recovers within minutes, and temporary local buffering is cheaper than dropping data. Increase consumers only when backend latency is low enough that parallel requests improve throughput rather than increasing contention. A practical review interval is weekly for high-volume services and monthly for lower-risk environments, with immediate review after traffic growth, a Collector upgrade, or an incident. A queue spanning several hours of traffic is a warning that the architecture may be absorbing a persistent mismatch between production and destination capacity. Retention should not be extended merely to avoid visible dropped-send counters; old telemetry has less diagnostic value than a clear, bounded failure mode.

Change architecture when the required recovery time exceeds what graceful shutdown and bounded retry can provide. Introduce a durable broker when separate scaling, replay, multiple downstream consumers, or stronger process-failure guarantees justify the added system. It is also appropriate when the same telemetry stream must feed several independent backends with different availability characteristics. Otherwise, the added broker may create a new bottleneck and another set of metrics to interpret. Since 1 October 2026, OpenTelemetry-related management and protocol tooling continues to evolve, including OpAMP-oriented collector management work, but that development does not automatically turn an in-memory Collector queue into durable storage. Verify supported features in the release being deployed rather than relying on announcements or examples from an earlier generation.

When loss is genuinely acceptable, document that policy and prioritize stable real-time metrics over exhaustive trace retention. For many observability systems, dropping old debug-level logs during an outage is preferable to an unbounded memory increase that threatens the entire Collector. The decision should be explicit: define which signals have priority, how long they may wait, and what percentage loss is tolerated during a five-, fifteen-, or sixty-minute incident. These thresholds are operational commitments, not defaults from OpenTelemetry. A queue configuration is mature when its failure behavior is understood, measured, and agreed upon by the people responsible for the service.

## Cost, Trade-offs, and the Practical Verdict

OpenTelemetry Collector software is generally free and open source, but queue memory, CPU, network egress, backend ingestion, and engineering time are not free. A 1 GiB memory allocation may be enough for a small Collector with modest retention, while 4 GiB or 8 GiB can be justified for higher volume or burst absorption; these are sizing examples, not recommended universal allocations. A durable broker adds storage and service costs, and SaaS observability vendors may price ingestion by data volume, span count, or retained records. Backends may also apply rate limits, making a larger local queue useful for short incidents but ineffective over a prolonged period. Compare the cost of a little dropped telemetry with the cost of scaling memory indefinitely.

The practical default is a bounded exporter queue, enabled retries with a finite elapsed time, a memory limiter with measured headroom, and a graceful shutdown period. Begin with documented default-like values, then adjust using real item sizes, p95 export latency, and outage tests rather than unverified claims about ideal consumer counts. Use an external broker only when the reliability requirement exceeds the Collector’s in-memory model. The result is not perfect zero-loss delivery, which no in-memory queue can promise, but controlled behavior with visible drops, predictable memory use, and enough resilience to handle ordinary traffic spikes. That is the right goal for OpenTelemetry queue configuration: not making every packet survive at any cost, but choosing and testing which telemetry can wait, for how long, and at what expense.

## Quick answers

### What is the default OpenTelemetry Collector exporter queue size?

Many exporter implementations use a default queue size of 2,048 items, but defaults vary by Collector distribution and version. Confirm the behavior in the documentation for your exact release, because the number of items also differs greatly in memory cost depending on payload size.

### Does an OpenTelemetry sending queue survive a Collector restart?

Standard exporter sending queues are generally held in memory and should not be assumed to survive a crash, forced termination, or restart. Use graceful shutdown and a durable external broker when telemetry must survive process failure or extended outages.

### Should I set num_consumers higher than 10?

Not automatically. Values above 10 may help some high-latency workloads, but backend rate limits and connection costs can make additional consumers worse. Test 2, 4, 8, and 10 consumers against real export duration, throughput, memory, and failure metrics.

### Is the OpenTelemetry memory limiter the same as an exporter queue?

No. A memory limiter controls Collector-wide memory pressure and can apply backpressure, while an exporter queue buffers pending telemetry for a particular destination. They solve different problems and are commonly configured together with memory headroom.

### How long should OpenTelemetry retry failed exports?

Choose a finite window based on the incident recovery objective rather than copying an arbitrary value. A five-minute retry window may suit short backend disruptions, while regulated or audit-oriented workloads may justify a durable broker instead of relying on longer in-memory retries.

Canonical: https://aitutorialmaker.com/knowledge/how_do_you_configure_opentelemetry_collector_queues_without_losing_telemetry.php
Markdown: https://aitutorialmaker.com/knowledge/how_do_you_configure_opentelemetry_collector_queues_without_losing_telemetry.php/index.md
