Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished Updated

KEDA ties Metrics Server readiness to the gRPC link behind every HPA query

KEDA merged fix #8226 so its Metrics Server reports NotReady when its gRPC link to the keda-operator is down. Before the change, the pod could stay Ready while the HPA behind it stopped getting any scale signal.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying KEDA ties Metrics Server readiness to the gRPC link behind every HPA query
Generated illustration
Metrics Server showed Ready while HPA scaling went silent How a Metrics Server that reports Ready while its gRPC link to the keda-operator is down affects each part of a KEDA cluster, before and after fix #8226.

HPA: loses scale signals while the replica looks healthy. Metrics Server replica: stays Ready with the operator link down; queries fail or hang. Debugging: pods look ready, hiding the cause. Builds with #8226: NotReady flags the gRPC link. Naive fix: can deadlock recovery.

Metrics Server showed Ready while HPA scaling went silent
WhoHowKindClaim
Kubernetes HPASilently stops receiving scale signals while the Metrics Server replica still looks healthyexposure2
Metrics Server replicaStayed Ready with the operator link down, stayed in the APIService endpoints, and every metric query failed or hungcontradiction6
Readiness-first debuggingThe first check is whether pods are ready, and they are, so the search points away from the real causeexposure18
Clusters on builds with #8226A NotReady Metrics Server pod indicates an unusable operator gRPC connectioncapability21
Naive version of the fixCan deadlock recovery: an Idle connection waits for an RPC, but RPCs resume only once the probe reports Readyconstraint11

What happened

  • The old probe passed whenever the Metrics Server process was up and its HTTP server was listening, without looking at the operator connection.
  • With the link down, Kubernetes kept the replica in the APIService endpoints and every HPA metric query routed to it failed or hung.
  • A maintainer found that a NotReady replica gets no RPCs, so gRPC's default 30-minute idle timeout parks its connection where nothing wakes it.

Why it matters

  • exposure Clusters on KEDA builds without #8226 can lose autoscaling while every Metrics Server pod reports Ready, so alerting keyed to pod readiness will not fire.
  • decision On a build with the fix, a NotReady Metrics Server pod points straight at the operator link, giving a stalled HPA an obvious first place to look.
  • constraint Any service whose readiness depends on a gRPC client faces the same trap: a probe that only reads GetState() can hold a pod out of rotation once the connection idles.

The Metrics Server does no metric work of its own. In KEDA, the CNCF event-driven autoscaler, the keda-operator talks to scalers and computes metrics [13][4]. The Metrics Server is the custom metrics API server the HPA queries. It forwards each request to the operator over gRPC and relays the answer [4]. "That gRPC connection is the Metrics Server's entire reason to exist," wrote the developer who hit the bug and submitted the first version of the fix, in a post on dev.to [14][17].

So a probe that reported Ready while that link was down was checking the wrong thing [1]. According to the post, the HPA stops receiving scale signals in that state while the replica still looks healthy [2]. "A crash is loud," the author wrote [15]. This failure sends debugging the wrong way, the post says, because the first thing anyone checks is whether the pods are ready, and they are [18].

Version one used state gRPC already tracks. grpc.ClientConn.GetState() returns Idle, Connecting, Ready, TransientFailure or Shutdown [7]. It reports state and never starts a connection, so a probe can call it with no side effects [7]. The patch added an accessor on the metrics-service client and a readyz check in the adapter that failed unless the state was Ready [8]. When the operator link dropped, Kubernetes pulled the replica from the APIService endpoints and the HPA stopped sending it queries [8].

The review comment is the more instructive part of this change. A maintainer spotted a failure that exists only because version one worked [9]. The post sets out the cycle like this:

1. The operator connection goes bad and the probe reports NotReady [11]. 2. Kubernetes removes the replica from the APIService endpoints, and no RPCs flow through that client [9]. 3. With no RPCs, gRPC parks the connection in Idle. GRPC_CLIENT_IDLE_TIMEOUT_MS defaults to 30 minutes [10]. 4. The Idle connection waits for an RPC. RPCs come back only after the probe reports Ready, and Ready needs the connection awake [11].

"An idle gRPC connection does not reconnect on its own. It waits for someone to use it," the author wrote [16]. The post says a replica caught this way can stay NotReady forever [11].

Kubernetes did exactly what the probe told it to in both bugs [6][9]. According to the post, the merged change has the readiness check nudge an Idle connection to reconnect [12]. The check gives up the property GetState() was picked for, because it can now start a connection as well as observe one [19]. We think that is the right tradeoff for this component. Once the replica is out of the endpoints, no RPC will touch that client, and the readiness check is the one code path left that can [20].

What to watch

  • The first KEDA release whose notes list #8226; from that version on, Metrics Server readiness reflects the operator link.
  • Whether the reconnect nudge inside the probe adds reconnect load on the keda-operator when many Metrics Server replicas probe an unreachable operator at once.
  • Whether other custom metrics adapters that proxy to a backend over gRPC copy the check-and-nudge readiness pattern.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence55
Adoption
Insufficient
Hype gap+5
Incentives35
Confidence55
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    KEDA's Metrics Server reported Ready even when its gRPC connection to the keda-operator was down.

    ReportedSupportedSource: dev.to post by the developer who submitted the fixView cited source
  2. [2]

    In that state the HPA silently stops receiving scale signals, while the replica still looks healthy.

    ReportedSupportedSource: dev.to postView cited source
  3. [3]

    The fix merged as kedacore/keda #8226; it makes the readiness probe observe the real gRPC connection state and report NotReady when the connection is unusable.

    ReportedSupportedSource: dev.to postView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · October 10, 2026

    Your Readiness Probe Is Lying to You

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories