A production Kubernetes cluster serving API traffic showed p95 latency stuck at 200ms despite adequate CPU and memory headroom. Packet captures revealed excessive east-west hairpinning and CNI conntrack pressure under burst load.
We tightened NetworkPolicies to reduce unnecessary cross-namespace flows, switched to a CNI with eBPF-based routing, and tuned kube-proxy mode on affected node pools. HPA thresholds were recalibrated after observing new baseline metrics.
Latency dropped from 200ms to 15ms at p95 within 48 hours. The fix was operational tuning — not throwing more nodes at the problem.