Skip to content
Technical Article Gabriel Illés, Senior DevOps Engineer

Enhance OpenTelemetry gRPC With a Consistent Hash Load Balancer

Architecture diagram: OpenTelemetry collector deployed as an agent on remote servers, sending telemetry through a gateway on the Kubernetes cluster into central storage

The use case

The OpenTelemetry collector (OTel collector) is deployed as an agent alongside the application on remote servers. It sends telemetry data — logs, traces, metrics — from the application and the host into central storage through a gateway deployed on the Kubernetes cluster.

The OTel collector is deployed using the OpenTelemetry operator Helm chart, with Kubernetes HPA scaling replicas based on CPU load. The traffic is routed through a headless service, because the standard Kubernetes service is not a good fit for gRPC. But with this setup, there is no load balancing on the Kubernetes side.

This lack of load balancing — combined with the OTel agents configured to send data in batches — causes the data from the same remote host to be forwarded randomly across the OTel collector gateway replicas. The data is written as many times as there are replicas into the storage, because of the different label values holding the identity of the OTel replica. This drastically increases storage usage, and the queries then have to be aggregated.

Let's show it in an example. Take one of the OTel agent metrics called otelcol_process_uptime, which has a label added by the OTel gateway called otelcol_replica, holding the name of the replica. The OTel gateway has four replicas; let's query the metric using PromQL on the storage side:

avg by (otelcol_replica)(otelcol_process_uptime{hostname="xxxxxx"})

{otelcol_replica="opentelemetry-collector-5fc9f8g5sj5"} 2502046.749352578
{otelcol_replica="opentelemetry-collector-5fc9f8pfmvh"} 2502096.748897170
{otelcol_replica="opentelemetry-collector-5fc9f8rzkh4"} 2502156.749325255
{otelcol_replica="opentelemetry-collector-5fc9f8xj95v"} 2502136.749453457

As demonstrated, the data coming from the remote host is written four times into the storage.

So the solution to this problem is a load balancing mechanism that provides consistency in routing data from the same remote source through the same OTel collector replica. And that's where the envoy-proxy is a perfect candidate, offering load balancers based on consistent hashing.

The solution

The envoy-proxy is deployed with two replicas and a headless service between the ingress and the OTel collector gateway.

It is configured with a ring-hash load balancer based on the X-Forwarded-For HTTP header, with HTTP2 enabled for the upstream clusters.

...
route:
  cluster: "opentelemetry-collector-cluster"
  hash_policy:
    - header:
        header_name: x-forwarded-for
...
clusters:
- name: opentelemetry-collector-cluster
  connect_timeout: 0.25s
  type: STRICT_DNS
  dns_lookup_family: V4_ONLY
  lb_policy: RING_HASH
  http2_protocol_options: {}
...
Architecture diagram: an envoy-proxy with a ring-hash load balancer placed between the ingress and the OTel collector gateway

This configuration ensures that the data from the same source IP will flow through the same OTel gateway replica while it exists. With this consistent route, only one copy of the data is written into storage from the remote host.

In case a replica fails, the envoy-proxy redirects the data flow to the next member of the hash ring. So for a short period, two copies of the data will exist in storage, due to the changed value of the label holding the identity of the OTel collector replica.

Conclusion

Consider a high-load environment where the number of OTel gateway replicas could be scaled to a high number. How much storage capacity could be saved with a reliable, deduplicated data flow from remote sources?

Where We Apply This

Reliable, cost-efficient observability is part of how we build and operate platforms for clients in regulated, high-stakes environments — where every replica, every label, and every byte written into storage adds up. Explore our infrastructure services or meet the team behind Grow2FIT.

Working on observability at scale? Let's talk.

From OpenTelemetry pipelines to full platform builds, we help teams see what their systems are really doing — without paying for it many times over in storage. Tell us what you're working on.

Schedule a call with us