The cnpg admission webhook's controller pod passes readinessProbe as soon as the local HTTPS server binds to :9443. The HelmRelease marks itself Ready right after that via helm install --wait. But the Service cnpg-webhook-service needs its EndpointSlice populated and the data plane (kube-proxy / cilium) programmed before kube-apiserver can reach the webhook through the Service ClusterIP. That gap is short but not zero, and any HelmRelease that depends on postgres-operator (cozy-keycloak, tenant Postgres apps) can fire its own install inside the window and hit Internal error occurred: failed calling webhook "mcluster.cnpg.io": failed to call webhook: Post "https://cnpg-webhook-service.cozy-postgres-operator.svc:443/...": dial tcp <svc-ip>:443: connect: connection refused which fails the install of the downstream release. Seen on cozystack/cozystack#2470 E2E run 24862782568. Add a post-install,post-upgrade Helm hook Job that blocks the release from reporting Ready until the webhook answers /readyz through the apiserver service proxy. Apiserver proxy routes the call over the same Service IP → EndpointSlice → pod path the admission webhook uses, so once it responds, the webhook admission path is also working. RBAC is minimal: a dedicated ServiceAccount with a ClusterRole that only grants get on services/proxy scoped to https:cnpg-webhook-service:webhook-server. The Job times out after 120s with 60 attempts at 2s intervals — longer than any data-plane programming delay seen on E2E, but bounded. Assisted-By: Claude <noreply@anthropic.com> Signed-off-by: Aleksei Sviridkin <f@lex.la>
14 lines
336 B
Makefile
14 lines
336 B
Makefile
export NAME=postgres-operator
|
|
export NAMESPACE=cozy-$(NAME)
|
|
|
|
include ../../../hack/package.mk
|
|
|
|
test:
|
|
helm unittest .
|
|
|
|
update:
|
|
rm -rf charts
|
|
helm repo add cnpg https://cloudnative-pg.github.io/charts
|
|
helm repo update cnpg
|
|
helm pull cnpg/cloudnative-pg --untar --untardir charts --version 0.26.1
|
|
rm -rf charts/cloudnative-pg/charts
|