# How Do You Tune OpenTelemetry Collector Queues Without Causing Data Loss?

aitutorialmaker.com · September 28, 2026

> Direct Answer OpenTelemetry Collector queue tuning starts with identifying which queue is actually under pressure: the receiver-side memory queue, the...

## Direct Answer

OpenTelemetry Collector queue tuning starts with identifying which queue is actually under pressure: the receiver-side memory queue, the exporter sending queue, or the batch processor queue. The most useful baseline is a sending queue of 2,000 to 10,000 requests, a batch size of 512 to 2,048 requests, and a 200 to 5,000 millisecond timeout, but those values are only a starting point. As of September 2026, the correct settings depend more on payload size, accepted spans per second, backend rate limits, and the collector’s available memory than on a universal formula. Queue tuning cannot create backend capacity, and increasing queue sizes can delay telemetry long enough to make a healthy pipeline look stalled. The aim is controlled short-term buffering, not unlimited retention. A well-tuned collector absorbs normal bursts, applies backpressure when an exporter cannot keep up, and preserves the oldest useful data during a brief outage without exhausting the host.

**Also worth reading:** [How Should You Set LLM Observability Alerts Without Causing Alert Fatigue?](https://aitutorialmaker.com/knowledge/how_should_you_set_llm_observability_alerts_without_causing_alert_fatigue.php) · [How Do You Test Authorization Controls on MCP Tools Without Exposing Production Data?](https://aitutorialmaker.com/knowledge/how_do_you_test_authorization_controls_on_mcp_tools_without_exposing_production_data.php) · [How Do OpenTelemetry GenAI Conventions Work for LLM and Agent Observability in 2026?](https://aitutorialmaker.com/knowledge/how_do_opentelemetry_genai_conventions_work_for_llm_and_agent_observability_in_2026.php)

## Queue Architecture and Failure Boundaries

The OpenTelemetry Collector normally places a memory limiter early in the pipeline, receives data through receiver components, processes it, and sends it through exporter components. The exporter helper commonly creates a sending queue, while the batch processor can have a separate queue if configured to do so. These queues solve different problems. The sending queue decouples receiver throughput from destination availability, whereas the batch processor groups records for efficient export. The memory limiter protects the process by refusing or degrading input before the Go runtime and operating system begin swapping or the container reaches an out-of-memory termination. These controls are related, but changing one does not automatically tune the others.

Backpressure is an intentional outcome, not necessarily a failure. If a destination processes 10,000 spans per second and a receiver temporarily admits 15,000, an exporter queue can retain the excess while new data waits. A bounded queue prevents this waiting period from growing without limit. However, memory queues are stored in RAM, and a queue advertised as holding 100,000 spans might occupy several hundred megabytes or more if each span averages 5 to 10 KiB. Reliable sizing therefore requires multiplication rather than reliance on request count. Queue persistence can protect against a collector restart, but it introduces disk requirements, replay behavior, and data-age considerations that a memory queue does not have.

## Core Configuration and Practical Tuning

Start by measuring the collector for at least one normal peak period. Record accepted data points per second, refused data points, queue utilization, exporter request duration, process resident memory, CPU saturation, and backend throttling. Then change one variable at a time and repeat the same workload. The usual first adjustment is the sending queue in the exporter helper: raise it enough to cover a measured destination outage while keeping its worst-case memory budget acceptable. A common initial range is 5,000 to 20,000 queued requests, but high-cardinality telemetry can make that unsafe on a 512 MiB collector. For a 1 GiB collector, 5,000 requests with an estimated 10 KiB average payload would already represent roughly 50 MiB before copies and batch overhead.

Batch settings should be tuned around destination request limits and latency, not just throughput. Increasing send_batch_size from 512 toward 2,048 can reduce request count and improve efficiency when the backend supports large batches. Increasing send_batch_max_size without raising the normal batch size may have little effect, because the processor rarely reaches the maximum. Queue capacity and batch size also interact: larger batches can improve exporter throughput but may exceed a gateway’s request-body limit or increase latency. A safe process is to begin with 1,024 sends, 1,024 maximum sends, and a 1-second timeout, then double the batch only if profiling shows excess request overhead. Validate the result with a controlled outage rather than assuming a larger queue is always better.

## Memory, CPU, and Container Limits

A Collector memory limit should be set for the container, not copied blindly from host memory. The memory limiter’s spike limit, data rate, and check interval determine how quickly it reacts, while the process must remain below the container’s cgroup or orchestration limit. Leave headroom for Go runtime behavior, goroutines, extensions, receivers, processors, and exporter serialization. As a conservative starting point, set the Collector’s memory limit to roughly 60 to 75 percent of the container allocation; for example, a 2 GiB container might begin near 1.5 GiB while retaining 0.5 GiB of headroom. That percentage is operational guidance, not a vendor guarantee, and bursty environments may need a wider margin.

A full queue can increase CPU use as producers continue to marshal data and the collector begins rejecting inputs. High rejection counts should therefore be investigated together with memory and destination latency. If the exporter times out repeatedly, retaining more records may only fill a larger RAM queue before the same outage ends. Persistent queueing should be considered when restart survival is worth additional disk I/O and possible stale-data export. It is especially relevant for security, billing, or audit telemetry whose delayed delivery remains useful. For approximate traces, dropping old observations during a prolonged backend outage may be acceptable. The retention decision belongs to the data owner and should be stated in the service’s telemetry policy.

## Comparing Memory and Persistent Queues

The most important queue decision is usually whether a burst should be buffered only in RAM or written to disk. Persistent storage is not automatically more reliable, because an undersized disk, slow volume, or overloaded backend can turn the queue into a source of stale data and duplicate replay. It also requires operators to monitor disk capacity and understand the queue’s cleanup and corruption behavior. Use the following comparison to frame the choice rather than treating persistent queues as a default recommendation.

| Feature | Memory queue | Persistent queue |
| --- | --- | --- |
| Survival across Collector restart | Lost or discarded unless another delivery layer exists | Can survive supported process restarts |
| Typical resource cost | RAM; approximately payload bytes multiplied by queue entries and in-flight copies | Local disk plus RAM for active batches and indexing |
| Best outage duration | Seconds to a few minutes, subject to RAM budget | Minutes to hours when disk and retention policy allow it |
| Main operational risk | Out-of-memory termination, pressure, or data loss | Disk exhaustion, slow replay, stale telemetry, or duplicates |
| Good starting use | Stateless, horizontally scaled telemetry pipelines | Audit, security, billing, or single-collector ingestion |

## Practical Procedure for Real Workloads
Begin with a representative test that preserves the production payload distribution. Synthetic tests containing only small spans can make a queue appear far larger than it is. Measure the mean and 95th-percentile encoded size, then estimate worst-case memory as queue capacity multiplied by the high-percentile item size. Include the batch currently in flight and accounting for the possibility that more than one serialized copy exists. If the result exceeds the collector’s safe memory budget, reduce queue capacity, use memory-limited batching, add a second replica, or improve the backend before raising limits again. A 20,000-entry queue is not a useful target if 20,000 worst-case payloads require 1 GiB.

Next, simulate a backend slowdown rather than only a total outage. Reduce destination capacity to 50 percent of normal for five minutes, then return it to 100 percent and observe recovery. Target a queue that covers the documented restart or incident window with at least 20 to 30 percent unused capacity at the beginning of the event. This margin absorbs a faster-than-expected burst and avoids starting at the hard limit. After recovery, watch the time required to drain the queue and verify that no new errors appear merely because the backlog is large. If a five-minute slowdown consumes 400,000 items, a 20,000-item queue is clearly inadequate unless dropping telemetry is intended; the remedy may require a larger persistent queue, backpressure, or more export capacity.

## Common Mistakes and Alternatives

The first common mistake is maximizing queue size because storage is initially cheap. A larger queue lengthens the time telemetry remains in the collector, delays alerts and debugging context, and can make batch behavior harder to predict. The second is treating rejected requests as an exporter problem without checking receiver-level backpressure. The third is setting a memory limiter after processors that already create large allocations; placement matters because early protection generally gives the collector more opportunity to regulate load. The fourth is assuming Kubernetes horizontal scaling alone replaces local buffering. Replicas add destination concurrency, but they do not preserve telemetry assigned to a failed pod unless acknowledgments and durable routing are designed accordingly.

Alternatives include enabling load balancing across healthy exporters, adding queue-aware routing, using a message broker, or separating lossy telemetry from records with strict delivery expectations. A message broker adds infrastructure and operating cost but can provide a durable boundary between collection and export. Direct export is simpler and cheaper when the backend is highly available and short outages are acceptable. For a typical three-replica Collector deployment, low-cost direct export is often the rational first choice; a managed Kafka-compatible service may cost more but can support longer buffering and consumer replay. Cost should include engineering time, observability, network transfer, broker retention, and the risk of exporting stale events, not only the monthly infrastructure invoice.

## When to Act and What Acceptance Criteria Should Mean

Act on queue tuning when evidence appears, such as sustained sending-queue utilization above roughly 70 to 80 percent, growing batch timeout, frequent exporter errors, rising refused data points, or recovery delayed by an oversized backlog. These are warning thresholds rather than universal standards: a queue that reaches 80 percent briefly during a normal burst may be appropriately sized. Do not act solely because a dashboard metric is red; correlate it with memory, backend limits, and the business requirement for loss. If data loss is unacceptable, first verify the end-to-end acknowledgment model because no local setting can guarantee exactly-once delivery across network and backend boundaries.

Define acceptance criteria before deployment. A useful test states the peak accepted rate, maximum average and 95th-percentile payload size, outage duration, maximum memory, acceptable refusal rate, and maximum telemetry age after recovery. For example, a team might require a five-minute backend interruption, less than 1 percent intentional refusal, no pod termination, and backlog drain within ten minutes after recovery. These targets must be adapted to workload and contractual requirements. A Collector that passes this test under 10,000 spans per second has not been proven at 1,000,000 spans per second, so production load testing remains necessary. Queue tuning is successful when failure is bounded, visible, and operationally understood rather than when every metric reaches an idealized maximum.

## Cost, Versions, and Configuration Validation

OpenTelemetry Collector distributions are generally open-source software and do not require a per-queue license fee. The direct monetary cost is the compute and memory used to buffer data, plus any managed load-balancer, gateway, or message-broker service added for resilience. Queueing can increase memory cost by 20 to several hundred percent compared with direct export, depending on burst size and retention duration; it can also reduce costs by batching more efficiently and limiting failed requests. Configuration semantics can vary by distribution, component, and major release, so validate examples against the documentation matching the deployed binary. In September 2026, treat vendor examples as inputs to a test plan, not proof that the same defaults are appropriate for another environment.

Use the Collector’s effective configuration output, component logs, internal telemetry, and backend receipt counts during validation. Confirm that the receiver, memory limiter, batch processor, and exporter helper are placed in the intended order, and ensure the limit is measured against the actual container allocation. Test a rolling restart as well as a backend outage, because a memory queue will not protect unacknowledged data through process termination. Finally, document queue capacity, expected worst-case bytes, retention time, refill rate, and an operator response for each threshold. The provided IANA port reference is useful for checking that network paths use registered or organization-approved ports, but it does not determine queue sizing; application throughput, payload size, and destination behavior are the controlling variables.

## Quick answers

### What is the best queue size for an OpenTelemetry Collector?

There is no universal best size. Start around 2,000 to 10,000 sending-queue items, then calculate worst-case memory from the measured 95th-percentile payload size and test against expected backend outages.

### Does a larger OpenTelemetry queue prevent data loss?

Only within its capacity and lifetime. Memory queues protect against brief exporter slowdowns, while persistent queues may survive supported restarts, but neither guarantees delivery across every failure mode.

### Should the OpenTelemetry memory limiter be set to the container limit?

Usually not. The memory limiter should normally remain below the container allocation, often around 60 to 75 percent as a conservative starting point, so runtime and temporary allocations have headroom.

### Is a persistent queue always better than a memory queue?

No. Persistent queues add disk I/O, capacity management, replay age, and duplicate considerations. They are most useful when data must survive Collector restarts and stale buffered telemetry still has business value.

### How often should queue settings be tuned?

Review them after meaningful workload, payload, backend, or container changes and whenever peaks produce sustained utilization above roughly 70 to 80 percent. Repeat controlled slow-down and restart tests after each material adjustment.

Canonical: https://aitutorialmaker.com/knowledge/how_do_you_tune_opentelemetry_collector_queues_without_causing_data_loss.php
Markdown: https://aitutorialmaker.com/knowledge/how_do_you_tune_opentelemetry_collector_queues_without_causing_data_loss.php/index.md
