Keycloak on Kubernetes with operators, step by step
Four years after our Keycloak.X on Kubernetes article, the state of the art has changed completely. Keycloak has an official operator, PostgreSQL has a mature one, and the Gateway API is taking over from Ingress. This article deploys a typical Keycloak platform on Kubernetes, step by step, with operators only and plain kubectl:
- CloudNativePG runs a PostgreSQL cluster: one primary, two standbys, synchronous replication;
- the Keycloak Operator runs two clustered Keycloak pods and imports a realm;
- Envoy Gateway exposes Keycloak, and protects applications that know nothing about OpenID Connect: it signs users in, checks their tokens and roles, and passes their identity to legacy applications in HTTP headers, from the token or from the UserInfo endpoint.
Nothing here depends on a particular Kubernetes distribution. Everything was deployed and tested end to end on the Kubernetes cluster of Docker Desktop: 48 automated checks, plus a crash of the PostgreSQL primary and a restart of every Keycloak pod while users keep signing in, and a rotation of the signing key.
Repository: https://github.com/please-openit/keycloak-kubernetes
This article and its repository are a reference to learn from and adapt. They come without any support and without any warranty: review them and test them in your own environment before relying on them.
If you only need a Keycloak, you probably do not need Kubernetes. Running Keycloak on Kubernetes is not trivial: three operators, a PostgreSQL cluster, network policies, certificates, upgrades, backups… For simplified deployments, a managed Keycloak such as our Keycloak as a Service with Clever Cloud takes all of that off your hands.
Kubernetes is a platform decision, not a Keycloak decision. You do not deploy a Kubernetes cluster just for Keycloak. You deploy Keycloak into an existing cluster when your platform already runs on Kubernetes, and your teams already know how to operate it: then Keycloak benefits from the same tooling, the same deployment pipeline, the same monitoring. The choice must make sense at the scale of the whole infrastructure.
Keycloak runs very well outside Kubernetes. With Docker (see Keycloak on Dokku) or directly on virtual or physical machines, including as a cluster: Keycloak nodes discover each other through the database and share their caches, wherever they run.
Autoscaling is rarely a reason. Adding Keycloak nodes on the fly only pays off for very specific traffic peaks, like Black Friday or the opening of ticket sales for a big event. In practice, very few platforms have them: a cluster sized for its peak load, with a node of margin, is simpler and more predictable.
| 2022 | 2026 |
|---|---|
| Keycloak.X, a preview of the Quarkus distribution | Keycloak 26.8, Quarkus only since Keycloak 20 |
| No operator for Keycloak.X: hand-written Deployment | Official Keycloak Operator: Keycloak and KeycloakRealmImport resources, rolling updates |
| Cluster discovery through a headless Service and DNS | Discovery through the database (jdbc-ping), nothing to configure |
| Sessions lost when all the nodes restart | Persistent user sessions, stored in the database |
| A database “left as an exercise” | PostgreSQL managed by CloudNativePG: failover, replication, backups |
| Ingress | Gateway API. The retirement of the Ingress NGINX controller was announced for March 2026 |
https://*.kube.localhost
|
Envoy Gateway: 2 Envoy proxies, TLS termination
______________________________|______________________________
| | |
keycloak.kube.localhost intranet.kube.localhost api.kube.localhost
keycloak-admin.kube.localhost OIDC + JWT + UserInfo + role JWT + role
| | | |
Keycloak x 2 intranet userinfo-authz echo-api
| (legacy app) (ext_authz)
PostgreSQL x 3 (CloudNativePG)
| Namespace | Content |
|---|---|
cnpg-system | CloudNativePG operator |
envoy-gateway-system | Envoy Gateway controller, Envoy proxies, the public Gateway |
keycloak | Keycloak Operator, Keycloak, its PostgreSQL cluster |
demo | Demo applications |
| Component | Version |
|---|---|
| Kubernetes (Docker Desktop 4.93, kind provisioning, 1 node) | 1.36.1 |
| CloudNativePG | 1.30.1 (PostgreSQL 18.6) |
| Keycloak and Keycloak Operator | 26.8.0 |
| Envoy Gateway | 1.9.2 |
The cluster needs three things, which every serious Kubernetes platform provides, and so does Docker Desktop:
LoadBalancerServices: the Envoy proxies are exposed through one. Cloud providers create a load balancer, on premises MetalLB or an equivalent does it. Docker Desktop publishes it onlocalhost.- A default StorageClass, for the PostgreSQL volumes. Pick one backed by SSDs in production.
- A network plugin that enforces NetworkPolicies. Some managed offerings need it to be enabled explicitly. Without it, the policies of this article are silently ignored, and so is part of the security.
On the workstation: kubectl (kustomize is built in), git, curl, jq and openssl. With Docker Desktop, enable Kubernetes in Settings → Kubernetes (kind provisioning, one node is enough) and give the virtual machine around 8 GB of memory.
All the names end with .kube.localhost: curl and Chromium-based browsers resolve *.localhost to 127.0.0.1 without any DNS setup.
git clone https://github.com/please-openit/keycloak-kubernetes.git
cd keycloak-kubernetes
Each step below shows its commands. ./deploy.sh runs them all in order.
# Namespaces of the platform.
#
# The "gateway-access: public" label allows HTTPRoutes of the namespace to attach
# to the HTTPS listener of the public Gateway (see 01-gateway/gateway.yaml).
#
# Pod Security Admission: "restricted" for the demo applications. The Keycloak
# Operator Deployment does not set a securityContext, so the keycloak namespace
# enforces "baseline" and only warns about "restricted".
apiVersion: v1
kind: Namespace
metadata:
name: keycloak
labels:
gateway-access: public
pod-security.kubernetes.io/enforce: baseline
pod-security.kubernetes.io/warn: restricted
---
apiVersion: v1
kind: Namespace
metadata:
name: demo
labels:
gateway-access: public
pod-security.kubernetes.io/enforce: restricted
kubectl apply -f 00-namespaces.yaml
The gateway-access: public label is how the Gateway decides which namespaces may publish routes on it. Pod Security Admission rejects non-compliant pods in demo. In keycloak, it only warns: the Keycloak Operator Deployment sets no securityContext, and you will see the warning when you install it.
The Keycloak documentation recommends installing its operator with the Operator Lifecycle Manager (OLM). OLM installs operators from catalogs, resolves their dependencies, and keeps them up to date through subscriptions. It is built into OpenShift. On any other distribution, OLM is one more component that you install and operate yourself.
What OLM changes for Keycloak:
- Automatic upgrades by default. A new operator version is installed as soon as it is published. The Keycloak Operator deploys the Keycloak version that matches its own: upgrading the operator upgrades Keycloak, including its database schema, and there is no going back without a database restore. The Keycloak documentation itself strongly recommends
installPlanApproval: Manual. - Upgrades through the OLM API (subscriptions, install plans to approve), not through your usual deployment pipeline.
Without OLM, operators are installed from versioned manifests. Versions only change when you change a line in a file: upgrades go through code review and your usual pipeline, GitOps tools such as Argo CD or Flux apply them, and you handle the custom resource definitions yourself. This article installs all three operators without OLM, with the official kubectl commands. It does not cover OpenShift specifics (OperatorHub, Routes, security context constraints).
# CloudNativePG, in the cnpg-system namespace
kubectl apply --server-side -f \
https://raw.githubusercontent.com/cloudnative-pg/cloudnative-pg/release-1.30/releases/cnpg-1.30.1.yaml
# Keycloak Operator, in the keycloak namespace (requires git)
kubectl apply -k 'github.com/keycloak/keycloak-k8s-resources/kubernetes?ref=26.8.0'
# Envoy Gateway, Gateway API definitions included, in the envoy-gateway-system namespace
kubectl apply --server-side -f \
https://github.com/envoyproxy/gateway/releases/download/v1.9.2/install.yaml
kubectl -n cnpg-system rollout status deployment/cnpg-controller-manager
kubectl -n keycloak rollout status deployment/keycloak-operator
kubectl -n envoy-gateway-system rollout status deployment/envoy-gateway
A few things worth knowing:
--server-sideis required for CloudNativePG and Envoy Gateway: their resource definitions are too large for a client-sidekubectl apply.- The Keycloak Operator installed this way only watches its own namespace: the
Keycloakresource has to live inkeycloak. A cluster-wide installation exists, still in preview. - Envoy Gateway’s
install.yamlincludes the Gateway API resource definitions. If your provider already manages them (some managed clusters do), follow the Envoy Gateway documentation for that case instead.
The Keycloak Operator can create an Ingress for Keycloak. We disable it: one Gateway API entry point serves Keycloak and the applications alike, with the same TLS configuration and the same security policies.
The HTTPS listener needs a certificate for *.kube.localhost. Locally, scripts/gen-certs.sh creates a small certificate authority with openssl and signs a wildcard certificate with it. In a real cluster, cert-manager or your own PKI does this job.
scripts/gen-certs.sh
kubectl -n envoy-gateway-system create secret tls kube-localhost-tls \
--cert=certs/tls.crt --key=certs/tls.key
# Public entry point of the platform: one GatewayClass handled by Envoy Gateway,
# one Gateway with an HTTP listener (redirect only) and an HTTPS listener.
apiVersion: gateway.networking.k8s.io/v1
kind: GatewayClass
metadata:
name: envoy
spec:
controllerName: gateway.envoyproxy.io/gatewayclass-controller
---
# Settings of the Envoy proxies deployed for the "public" Gateway.
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: EnvoyProxy
metadata:
name: public
namespace: envoy-gateway-system
spec:
# By default ext_authz runs before the OIDC (oauth2) and JWT filters. Moving it
# after jwt_authn gives the external authorization service the access token
# obtained by the OIDC filter, and the claims already validated by the JWT filter.
filterOrder:
- name: envoy.filters.http.ext_authz
after: envoy.filters.http.jwt_authn
provider:
type: Kubernetes
kubernetes:
envoyDeployment:
replicas: 2
container:
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
memory: 512Mi
envoyPDB:
minAvailable: 1
---
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: public
namespace: envoy-gateway-system
spec:
gatewayClassName: envoy
infrastructure:
parametersRef:
group: gateway.envoyproxy.io
kind: EnvoyProxy
name: public
listeners:
- name: http
protocol: HTTP
port: 80
hostname: "*.kube.localhost"
allowedRoutes:
namespaces:
from: Same
- name: https
protocol: HTTPS
port: 443
hostname: "*.kube.localhost"
tls:
mode: Terminate
certificateRefs:
- kind: Secret
name: kube-localhost-tls
allowedRoutes:
namespaces:
from: Selector
selector:
matchLabels:
gateway-access: public
---
# Plain HTTP only redirects to HTTPS.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: https-redirect
namespace: envoy-gateway-system
spec:
parentRefs:
- name: public
sectionName: http
rules:
- filters:
- type: RequestRedirect
requestRedirect:
scheme: https
statusCode: 301
---
# Headers that only the gateway may set are removed from every incoming request,
# before any filter runs: a client cannot spoof an identity header.
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: ClientTrafficPolicy
metadata:
name: public
namespace: envoy-gateway-system
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: public
headers:
earlyRequestHeaders:
remove:
- X-Remote-User
- X-Remote-Name
- X-Remote-Email
- X-Remote-Groups
- X-Remote-Employee-Id
- X-Remote-Department
- X-Client-Id
# Envoy is the first proxy here: client values are not trusted.
- X-Forwarded-For
- X-Forwarded-Host
- X-Forwarded-Port
- Forwarded
kubectl apply -f 01-gateway/
kubectl -n envoy-gateway-system wait gateway/public --for=condition=Programmed
EnvoyProxyruns two Envoy replicas with a PodDisruptionBudget. ItsfilterOrderis explained with the intranet.- The
httpslistener only accepts routes from namespaces labeledgateway-access: public. Thehttplistener only redirects to HTTPS. ClientTrafficPolicyremoves, from every incoming request and before anything else, the headers that only the gateway may set. Without it, a client could send its ownX-Remote-Userheader to an application that trusts it. We come back to this below.
The Gateway gets an address from the LoadBalancer Service. Docker Desktop also publishes it on localhost:
$ kubectl -n envoy-gateway-system get gateway
NAME CLASS ADDRESS PROGRAMMED AGE
public envoy 172.29.0.5 True 7m13s
# PostgreSQL cluster dedicated to Keycloak, managed by CloudNativePG:
# one primary and two standbys, with automated failover.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: keycloak-db
namespace: keycloak
spec:
instances: 3
# Pinned image: the PostgreSQL version changes only when this line changes.
imageName: ghcr.io/cloudnative-pg/postgresql:18.6-system-trixie
# Creates the "keycloak" database owned by the "keycloak" user. The operator
# stores the generated credentials in the "keycloak-db-app" Secret.
bootstrap:
initdb:
database: keycloak
owner: keycloak
postgresql:
parameters:
shared_buffers: 128MB
max_connections: "100"
# A commit waits for at least one standby: no committed transaction is
# lost on failover (RPO = 0).
synchronous:
method: any
number: 1
# Defaults made explicit. On a multi-node cluster, use "required" so that two
# instances never share a node; spread across zones with
# topologyKey: topology.kubernetes.io/zone.
affinity:
enablePodAntiAffinity: true
topologyKey: kubernetes.io/hostname
podAntiAffinityType: preferred
storage:
size: 2Gi
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
memory: 512Mi
- Three instances: one primary, two standbys that stream its changes. When the primary fails, CloudNativePG promotes the most up-to-date standby and repoints the
keycloak-db-rwService to it. - Quorum-based synchronous replication: a transaction is only committed once at least one standby has it, so a registration or a password change is already on a standby when the primary fails, and CloudNativePG promotes the most up-to-date standby. With three instances, losing one of them does not block writes.
- The credentials of the
keycloakuser are generated by the operator, in thekeycloak-db-appSecret. Keycloak reads them from there: no password is written anywhere. - Anti-affinity:
preferredplaces the instances on different nodes when it can, and still works on our single node. On a real cluster, setrequired: three instances on the same node are not highly available.
The database only accepts connections from its own namespace and from the CloudNativePG operator:
# Only the pods of the keycloak namespace (Keycloak, its import Jobs and the
# other PostgreSQL instances) and the CloudNativePG operator reach the database.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: keycloak-db
namespace: keycloak
spec:
podSelector:
matchLabels:
cnpg.io/cluster: keycloak-db
policyTypes:
- Ingress
ingress:
- from:
- podSelector: {}
ports:
- port: 5432
- port: 8000
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: cnpg-system
podSelector:
matchLabels:
app.kubernetes.io/name: cloudnative-pg
ports:
- port: 5432
- port: 8000
kubectl apply -f 02-database/
kubectl -n keycloak wait cluster/keycloak-db --for=condition=Ready --timeout=10m
$ kubectl -n keycloak get cluster keycloak-db
NAME AGE INSTANCES READY STATUS PRIMARY
keycloak-db 7m1s 3 3 Cluster in healthy state keycloak-db-2
Replication is not a backup: aDROPor a broken migration replicates too. In production, configure continuous backups to object storage with the Barman Cloud plugin, and test a point-in-time recovery before you need one. Back up the database before every Keycloak upgrade: database migrations cannot be rolled back.
Planned maintenance: CloudNativePG creates a PodDisruptionBudget that blocks draining the node of the primary. Move the primary first, with a switchover: kubectl cnpg promote keycloak-db <instance>, from the cnpg plugin. Deleting the primary pod instead does not fail over right away: the instance first waits up to 180 seconds (smartShutdownTimeout) for open connections to close, and Keycloak keeps its connections open.
# Keycloak cluster managed by the Keycloak Operator.
apiVersion: k8s.keycloak.org/v2beta1
kind: Keycloak
metadata:
name: keycloak
namespace: keycloak
spec:
instances: 2
# Database created by CloudNativePG: "-rw" Service of the cluster (always the
# primary) and the credentials Secret generated by the operator.
db:
vendor: postgres
host: keycloak-db-rw
port: 5432
database: keycloak
usernameSecret:
name: keycloak-db-app
key: username
passwordSecret:
name: keycloak-db-app
key: password
# Fixed-size pool. Instances x poolMaxSize must stay below max_connections.
poolInitialSize: 20
poolMinSize: 20
poolMaxSize: 20
# TLS terminates at the Gateway; Keycloak serves plain HTTP inside the cluster
# and the NetworkPolicy below restricts who can reach it.
http:
httpEnabled: true
# Public URL used in tokens (issuer) and redirects, and a separate URL for
# the administration console.
hostname:
hostname: https://keycloak.kube.localhost
admin: https://keycloak-admin.kube.localhost
# Envoy sets X-Forwarded-For and X-Forwarded-Proto.
proxy:
headers: xforwarded
# Exposure is handled by Gateway API HTTPRoutes (routes.yaml).
ingress:
enabled: false
# NetworkPolicy generated by the operator: the HTTP port only accepts traffic
# from the Envoy proxies and from the userinfo-authz service of the demo.
# The clustering port stays restricted to the Keycloak pods.
networkPolicy:
enabled: true
http:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: envoy-gateway-system
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: demo
podSelector:
matchLabels:
app.kubernetes.io/name: userinfo-authz
resources:
requests:
cpu: 500m
memory: 1Gi
limits:
memory: 1536Mi
# Rolling update when compatible (configuration change, patch release),
# recreate otherwise.
update:
strategy: Auto
additionalOptions:
- name: metrics-enabled
value: "true"
# Without spec.scheduling, the operator spreads the pods across zones and
# nodes. On a multi-node or multi-site cluster, state the rule explicitly:
# scheduling:
# affinity:
# podAntiAffinity:
# requiredDuringSchedulingIgnoredDuringExecution:
# - labelSelector:
# matchLabels:
# app: keycloak
# app.kubernetes.io/instance: keycloak
# topologyKey: kubernetes.io/hostname
The key settings:
db: the-rwService always points to the primary, failover included. The connection pool has a fixed size. Make sureinstances × poolMaxSizestays below PostgreSQL’smax_connections: the defaults (100 connections per Keycloak pod, 100 for PostgreSQL) do not fit together.http.httpEnabled: TLS terminates at the Gateway, and Keycloak serves plain HTTP inside the cluster. This is the common setup. If your security policy requires encryption up to the pod, give Keycloak a certificate (http.tlsSecret) and Envoy aBackendTLSPolicy.hostname: the public URL. It is the issuer of every token, whatever URL a request comes from. Applications inside the cluster can call Keycloak through its internal Service, the tokens still carry the public issuer.admingives the administration console its own hostname.proxy.headers: xforwarded: Keycloak trustsX-Forwarded-*headers. That is only safe because Envoy sets them, and because nobody else can reach Keycloak: hence the NetworkPolicy.networkPolicy: the operator generates it. The HTTP port only accepts the Envoy proxies and theuserinfo-authzservice of the demo. The clustering ports only accept the other Keycloak pods.resources: the operator defaults are 1700 MiB requested and 2 GiB limit. We lowered them to fit on a laptop. The heap is sized as a percentage of the limit. Size production with the Keycloak sizing guide.update.strategy: Auto: configuration changes and patch releases roll out one pod at a time without downtime. Other changes recreate the pods.
We add a PodDisruptionBudget, which the operator does not create:
# Voluntary disruptions (node drain, cluster upgrade) never stop both
# Keycloak pods at the same time.
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: keycloak
namespace: keycloak
spec:
maxUnavailable: 1
selector:
matchLabels:
app: keycloak
app.kubernetes.io/instance: keycloak
kubectl apply -f 03-keycloak/
kubectl -n keycloak wait keycloak/keycloak --for=condition=Ready --timeout=10m
The two pods form a cluster on their own. They find each other through the database (jdbc-ping, the default), with no discovery Service to configure:
$ kubectl -n keycloak logs keycloak-0 | grep ISPN000094
... ISPN000094: Received new cluster view for channel ISPN: [keycloak-1(...)|5] (2) [keycloak-1(...), keycloak-0(...)]
The operator stores a temporary administrator, temp-admin, in the keycloak-initial-admin Secret. Use it to create your own administrators, then delete it.
The stock image runs a build step at each start before Keycloak itself starts. On a laptop, two pods starting together, plus three PostgreSQL instances, are enough to starve the node: on our first attempt, health probes timed out until things settled down by themselves. In production, build an optimized image (kc.sh build, with your themes and extensions) and setstartOptimized: true.
Without spec.scheduling, the operator spreads the Keycloak pods across zones and nodes, with ScheduleAnyway constraints: a preference, not a rule. That is why it works on a single node. As soon as the cluster has several nodes, or spans several sites, define anti-affinity explicitly: two Keycloak pods on the same node, or three PostgreSQL instances in the same zone, do not survive the loss of that node or zone. Use requiredDuringScheduling... rules with kubernetes.io/hostname and topology.kubernetes.io/zone (see the comments in the manifests above), on both Keycloak and CloudNativePG. Spanning several sites calls for more than that: see Keycloak’s multi-site documentation. We did not test these settings, our cluster has a single node.
Keycloak’s reverse proxy guide lists the paths that must be public: /realms/, /resources/ and /.well-known/. The administration console, the metrics, the health checks and the master realm do not need to be.
# Public hostname: only the paths needed by applications and users.
# /admin, /metrics, /health and the master realm are not reachable here.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: keycloak-public
namespace: keycloak
spec:
parentRefs:
- name: public
namespace: envoy-gateway-system
sectionName: https
hostnames:
- keycloak.kube.localhost
rules:
- matches:
- path:
type: PathPrefix
value: /realms/master
filters:
- type: ExtensionRef
extensionRef:
group: gateway.envoyproxy.io
kind: HTTPRouteFilter
name: not-found
- matches:
- path:
type: PathPrefix
value: /realms
- path:
type: PathPrefix
value: /resources
- path:
type: PathPrefix
value: /.well-known
backendRefs:
- name: keycloak-service
port: 8080
---
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: HTTPRouteFilter
metadata:
name: not-found
namespace: keycloak
spec:
directResponse:
statusCode: 404
contentType: text/plain
body:
type: Inline
inline: "Not found\n"
---
# Administration hostname: every path, restricted to an address allowlist.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: keycloak-admin
namespace: keycloak
spec:
parentRefs:
- name: public
namespace: envoy-gateway-system
sectionName: https
hostnames:
- keycloak-admin.kube.localhost
rules:
- backendRefs:
- name: keycloak-service
port: 8080
---
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: SecurityPolicy
metadata:
name: keycloak-admin-allowlist
namespace: keycloak
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: keycloak-admin
authorization:
defaultAction: Deny
rules:
- name: admin-networks
action: Allow
principal:
# Replace with the networks of your administrators (VPN, bastion...).
# Here: private ranges, the source address seen behind Docker Desktop.
clientCIDRs:
- 10.0.0.0/8
- 172.16.0.0/12
- 192.168.0.0/16
- On
keycloak.kube.localhost, only the three public paths are routed. Envoy answers 404 for everything else. Themasterrealm gets an explicit 404 through an Envoy GatewayHTTPRouteFilter. keycloak-admin.kube.localhostroutes everything, but aSecurityPolicyonly lets through clients from an address allowlist.
The allowlist only makes sense if Envoy sees the real client address. On Docker Desktop, it sees the load balancer’s address, 172.29.0.5, hence the private ranges. In production, preserve the client address (externalTrafficPolicy: Local, or PROXY protocol behind a cloud load balancer), restrict it to your administrators’ networks, or, better still, publish the console on a second Gateway that is only reachable from the internal network.
The KeycloakRealmImport resource creates a realm through a Job. The secrets come from a Kubernetes Secret through placeholders, they never appear in the file:
scripts/create-secrets.sh # random passwords and client secrets
kubectl apply -f 04-realm/
kubectl -n keycloak wait keycloakrealmimport/demo --for=condition=Done
| Object | Content |
|---|---|
| Users | alice, member of staff; bob, member of contractors; HR attributes employee_id and department, declared in the user profile |
intranet client | Confidential client of the Envoy OIDC filter, PKCE required. Audience intranet. Groups and HR attributes only in the UserInfo response |
echo-api client | The API: no login flow, only two client roles, read and write |
reporting-job, inventory-sync, audit-job | Machine clients (client credentials): read only, read + write, and no role of the API |
KeycloakRealmImport only creates a realm: if the realm exists, nothing happens, and later changes to the resource are ignored. To manage a realm over time, look at keycloak-config-cli or the Terraform/OpenTofu provider. The Keycloak Operator 26.8 also offers KeycloakOIDCClient and KeycloakSAMLClient resources to manage clients from Kubernetes, still in preview. And for production, review which roles end up in your tokens: see our article on the “full scope allowed” switch.
This is where Envoy Gateway goes beyond exposing Keycloak. A SecurityPolicy attached to a route makes Envoy handle authentication and authorization, so the application behind it does not have to. We deploy three small Python scripts (standard library only), mounted from a ConfigMap into the official Python image: no image to build, no registry.
kubectl apply -k 05-demo/
The Deployments and Services are ordinary: no authentication code, no authentication setting.
intranet stands for the old application that every company has: it has no idea what OpenID Connect is. Like an application behind an Apache authentication module or a web access management agent, it reads the identity of the user from HTTP headers:
Everything happens in its route and in the SecurityPolicy attached to it:
# Legacy intranet application, protected by Envoy Gateway without any change:
# OIDC login, token validation, UserInfo attributes, role check, identity headers.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: intranet
spec:
parentRefs:
- name: public
namespace: envoy-gateway-system
sectionName: https
hostnames:
- intranet.kube.localhost
rules:
- backendRefs:
- name: intranet
port: 8080
filters:
# The filters of the SecurityPolicy use the access token; the
# application does not need it, it never receives it.
- type: RequestHeaderModifier
requestHeaderModifier:
remove:
- Authorization
---
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: SecurityPolicy
metadata:
name: intranet
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: intranet
# 1. Authentication: without a valid session, the browser is redirected to
# Keycloak (authorization code flow). Envoy keeps the tokens in encrypted
# cookies and refreshes them.
oidc:
provider:
issuer: https://keycloak.kube.localhost/realms/demo
# Called by the browser: public URL.
authorizationEndpoint: https://keycloak.kube.localhost/realms/demo/protocol/openid-connect/auth
endSessionEndpoint: https://keycloak.kube.localhost/realms/demo/protocol/openid-connect/logout
# Called by Envoy: internal URL.
tokenEndpoint: http://keycloak-service.keycloak.svc.cluster.local:8080/realms/demo/protocol/openid-connect/token
clientID: intranet
clientSecret:
name: intranet-oidc-client
redirectURL: https://intranet.kube.localhost/oauth2/callback
logoutPath: /logout
scopes:
- profile
- email
# Passes the access token to the next filters in the Authorization header.
forwardAccessToken: true
# 2. Validation of the access token and claims -> headers.
jwt:
providers:
- name: keycloak
issuer: https://keycloak.kube.localhost/realms/demo
audiences:
- intranet
remoteJWKS:
uri: http://keycloak-service.keycloak.svc.cluster.local:8080/realms/demo/protocol/openid-connect/certs
claimToHeaders:
- header: X-Remote-User
claim: preferred_username
- header: X-Remote-Name
claim: name
- header: X-Remote-Email
claim: email
# 3. Attributes that are not in the token: UserInfo -> headers.
extAuth:
http:
backendRefs:
- name: userinfo-authz
port: 8080
path: /authz
headersToBackend:
- X-Remote-Groups
- X-Remote-Employee-Id
- X-Remote-Department
# 4. Authorization: the "user" role of the intranet client is required.
authorization:
defaultAction: Deny
rules:
- name: intranet-users
action: Allow
principal:
jwt:
provider: keycloak
claims:
- name: resource_access.intranet.roles
valueType: StringArray
values:
- user
A request goes through four steps, in this order:
Browser -> Envoy oauth2 no session cookie: redirect to Keycloak (code flow + PKCE)
Browser -> Keycloak login form, redirect to /oauth2/callback?code=...
Browser -> Envoy oauth2 code -> tokens (internal URL), encrypted cookies, redirect
Browser -> Envoy oauth2 valid session: Authorization: Bearer <access token>
jwt_authn signature, issuer, audience; claims -> X-Remote-User/Name/Email
ext_authz userinfo-authz -> Keycloak UserInfo -> X-Remote-Groups/...
rbac role "user" of the intranet client required, otherwise 403
router Authorization header removed -> intranet
Envoy runs the authorization code flow, with PKCE, and keeps the tokens in encrypted cookies. It refreshes them when they expire, and /logout clears them and ends the Keycloak session. The browser reaches Keycloak through the public URL. Envoy exchanges the code through the internal URL: that call never leaves the cluster. Providing every endpoint explicitly also spares the Envoy Gateway controller from fetching the discovery document through the public URL.
forwardAccessToken passes the access token to the next filter, which checks its signature against Keycloak’s keys (fetched through the internal URL), its issuer and its audience. claimToHeaders then copies preferred_username, name and email into headers.
claimToHeaders only handles strings, numbers and booleans: arrays are not supported, and groups are an array. They will come from UserInfo.
Envoy keeps Keycloak’s public keys in cache for 5 minutes (cacheDurationofremoteJWKS). A token signed with a key that Envoy does not know yet is rejected with401 Jwks doesn't have key to match kid or alg from Jwt: we hit it when we recreated the realm, with brand new keys. To rotate keys, add the new key with a lower priority than the current one, so that Keycloak publishes it without signing with it, wait longer than the cache duration, then raise its priority. Our tests check both ways.
Envoy has no native support for the UserInfo endpoint, but it can call an external authorization service (ext_authz) for each request. Ours is a short Python script: it calls UserInfo with the access token and turns the claims into headers. Envoy copies the headers listed in headersToBackend into the request sent to the application.
By default, Envoy Gateway runs ext_authz before the OIDC filter: on that request, the service would not see any token yet. The filterOrder of the EnvoyProxy resource (step 3) moves it after jwt_authn.
Why the extra call?
- Smaller tokens. Envoy stores the tokens in cookies, and browsers limit a cookie to about 4 KB. The groups of a user can be counted in hundreds. Keeping them out of the tokens (
access.token.claim: falseon the mappers) keeps the cookies small. - Up-to-date attributes, read at each call (here, cached for 10 seconds).
- Revocation. UserInfo only answers while the Keycloak session exists. End a session in the administration console, and the next request is refused once the cache expires, instead of when the access token expires (5 minutes by default). Our tests check it.
The price: one call to Keycloak per request and per cache period. Size the cache accordingly.
authorization reads a claim validated by the JWT filter: only users with the user role of the intranet client get through. Alice gets it from the staff group. Bob is a contractor:
$ curl ... https://intranet.kube.localhost/ # after signing in as bob
RBAC: access denied # HTTP 403
Alice gets through, and here is what the application receives (/whoami dumps its headers):
{
"identity": {
"X-Remote-User": "alice",
"X-Remote-Name": "Alice Martin",
"X-Remote-Email": "alice@example.com",
"X-Remote-Groups": "staff",
"X-Remote-Employee-Id": "E-1001",
"X-Remote-Department": "Finance"
},
"authorization_header": false,
"x_forwarded_for": "172.29.0.5",
...
}
The first three headers come from the token, the last three from UserInfo. The application never sees the access token: the RequestHeaderModifier filter of the route removes it, after the security filters have used it.
An application that trusts headers is only safe if nobody but the gateway can set them. That takes two conditions, and both are tested:
- Envoy removes these headers from every incoming request (the
ClientTrafficPolicyof step 3), then sets them itself. A client that sendsX-Remote-User: admingets nowhere. Removing them early matters: a header that the policy of a route does not set would otherwise reach the application as the client sent it. - The application is only reachable through Envoy:
# The applications trust identity headers: they must only be reachable
# through the Envoy proxies. Everything else is denied.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-ingress
spec:
podSelector: {}
policyTypes:
- Ingress
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-from-envoy
spec:
podSelector: {}
policyTypes:
- Ingress
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: envoy-gateway-system
ports:
- port: 8080
From any other pod of the cluster, the application cannot be reached at all, so a forged header has no way in.
echo-api is a REST API, called by machines with a bearer token. No redirect here: without a valid token, the answer is 401. The SecurityPolicy checks the token, and the role required by the HTTP method:
# REST API protected by bearer tokens: no redirect, 401 without a valid token,
# 403 without the role required by the HTTP method.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: echo-api
spec:
parentRefs:
- name: public
namespace: envoy-gateway-system
sectionName: https
hostnames:
- api.kube.localhost
rules:
- backendRefs:
- name: echo-api
port: 8080
---
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: SecurityPolicy
metadata:
name: echo-api
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: echo-api
jwt:
providers:
- name: keycloak
issuer: https://keycloak.kube.localhost/realms/demo
# Keycloak adds "echo-api" to the audience of tokens that carry one
# of its client roles.
audiences:
- echo-api
remoteJWKS:
uri: http://keycloak-service.keycloak.svc.cluster.local:8080/realms/demo/protocol/openid-connect/certs
claimToHeaders:
- header: X-Client-Id
claim: azp
- header: X-Remote-User
claim: preferred_username
authorization:
defaultAction: Deny
rules:
- name: read
action: Allow
operation:
methods: ["GET", "HEAD"]
principal:
jwt:
provider: keycloak
claims:
- name: resource_access.echo-api.roles
valueType: StringArray
values: ["read"]
- name: write
action: Allow
operation:
methods: ["POST", "PUT", "PATCH", "DELETE"]
principal:
jwt:
provider: keycloak
claims:
- name: resource_access.echo-api.roles
valueType: StringArray
values: ["write"]
The audiences check relies on a Keycloak mechanism: when a token carries client roles of echo-api, Keycloak adds echo-api to its audience. A token issued for another application, such as the one of audit-job, is rejected with 403 Audiences in Jwt are not allowed.
$ curl https://api.kube.localhost/items
Jwt is missing # HTTP 401
$ TOKEN=$(curl -s -d grant_type=client_credentials -d client_id=reporting-job \
-d client_secret=... https://keycloak.kube.localhost/realms/demo/protocol/openid-connect/token | jq -r .access_token)
$ curl -H "Authorization: Bearer $TOKEN" https://api.kube.localhost/items
{
"caller": {
"client_id": "reporting-job",
"user": "service-account-reporting-job",
"authorization_header": true
},
"items": [ ... ]
}
$ curl -H "Authorization: Bearer $TOKEN" -d '{"name":"mouse"}' https://api.kube.localhost/items
RBAC: access denied # HTTP 403: reporting-job has no "write" role
Envoy Gateway can do more for legacy applications, which we did not test here: inject a fixed Authorization header towards a service that only knows HTTP Basic authentication (HTTPRouteFilter with credentialInjection), or accept both browsers and API clients on the same route (passThroughAuthHeader). Outside Kubernetes, our NGINX-based authentication proxy and oauth2-proxy fill the same role.
The repository is not just manifests: test.sh checks the whole deployment from outside the cluster, through the Gateway, as a user would, and from inside, for the network policies. It drives the OpenID Connect login with curl, posting the Keycloak login form.
$ ./test.sh
Platform
PASS CloudNativePG: 3 ready instances
PASS PostgreSQL: quorum synchronous replication
PASS Keycloak CR Ready
PASS Keycloak: 2 ready pods
PASS Keycloak: 2-member Infinispan cluster
PASS Gateway programmed
PASS Envoy: 2 proxy replicas
Keycloak exposure
PASS OIDC discovery: public issuer
PASS HTTP redirects to HTTPS
PASS Public hostname: /admin not exposed
PASS Public hostname: master realm not exposed
PASS Public hostname: /metrics not exposed
PASS Admin hostname: console available
API (bearer token)
PASS No token: 401
PASS Invalid token: 401
PASS Forged signature: 401
PASS Token issued for another audience (audit-job): 403
PASS reporting-job GET (role read): 200
PASS reporting-job POST (no role write): 403
PASS inventory-sync POST (role write): 201
PASS Header X-Client-Id from the azp claim, client value ignored
Intranet (OIDC login, legacy application)
PASS Anonymous request redirected to Keycloak
PASS alice signs in: 200
PASS X-Remote-User from the token
PASS X-Remote-Name from the token
PASS X-Remote-Email from the token
PASS X-Remote-Groups from UserInfo
PASS X-Remote-Employee-Id from UserInfo
PASS X-Remote-Department from UserInfo
PASS The application does not receive the access token
PASS Spoofed X-Remote-User replaced
PASS Spoofed X-Remote-Employee-Id replaced
PASS Identity header not set by this route (X-Client-Id) removed
PASS Spoofed X-Forwarded-For removed (seen: 172.29.0.5)
PASS bob (contractors, no intranet role) signs in: 403
PASS Logout redirects to the Keycloak end session endpoint
PASS Keycloak ends the session and returns to the intranet
PASS After logout, the login form is shown again (no single sign-on)
Session revocation (UserInfo check)
PASS alice signed in again
PASS Administrator ends alice's sessions
PASS Access refused although the access token has not expired: 401
PASS Refused by the UserInfo check (ext_authz)
NetworkPolicies (from inside the cluster)
PASS default -> intranet:8080 blocked
PASS default -> keycloak:8080 blocked
PASS default -> postgresql:5432 blocked
PASS envoy-gateway-system -> intranet:8080 open
PASS envoy-gateway-system -> keycloak:8080 open
PASS envoy-gateway-system -> postgresql:5432 blocked
48 passed, 0 failed
test-failover.sh breaks things while a user signs in to the intranet every two seconds:
$ ./test-failover.sh
PostgreSQL: crash of the primary
PASS A standby is promoted (keycloak-db-1 -> keycloak-db-2, 19s)
PASS Cluster back to 3 ready instances
26 sign-ins, 9 failed (12:47:20 to 12:47:37)
PASS Sign-ins succeed after the failover
Keycloak: restart of every pod, one at a time
PASS alice signs in before the restarts
32 sign-ins, 0 failed
PASS Sign-ins succeed after the restarts
PASS Session opened before the restarts still valid (UserInfo)
6 passed, 0 failed
- Crash of the PostgreSQL primary (pod deleted without grace period): a standby is promoted 19 seconds later. Sign-ins fail for about 17 seconds, then resume by themselves: Keycloak reconnects to the new primary through the same Service.
- Restart of both Keycloak pods, one after the other: not a single failed sign-in, and the session opened before the restarts is still valid afterwards, since user sessions are stored in the database.
test-key-rotation.sh checks the key rotation procedure, in about six minutes:
$ ./test-key-rotation.sh
Before the rotation
PASS API call with the current key: 200
Naive rotation: the new key signs at once
PASS New key used for signing
PASS Envoy rejects tokens signed with it: 401
PASS Key removed
PASS Previous key signs again, API call: 200
Safe rotation: publish first, sign later
PASS New key published in the JWKS
PASS Previous key still signs
waiting 310 seconds for the JWKS cache of Envoy to expire...
PASS API call (Envoy fetches the JWKS again): 200
PASS New key gets the highest priority
PASS New key used for signing
PASS Envoy accepts tokens signed with it: 200
11 passed, 0 failed
What these tests do not cover: anti-affinity on several nodes or sites, backups and restores, and upgrades.
./teardown.sh removes everything. It deletes the PostgreSQL cluster, and waits for its pods to stop, before deleting the namespace: an instance whose ServiceAccount is already gone can no longer archive its WAL, and waits up to 30 minutes (stopDelay) before stopping.
- Certificates and DNS: cert-manager, real names, and a separate, internal entry point for the administration console.
- Backups: the Barman Cloud plugin for CloudNativePG, restores tested regularly.
- Anti-affinity set to
required, across nodes and zones, for Keycloak and PostgreSQL. - An optimized Keycloak image with your extensions and themes, and sizing based on your load.
- Secrets: generated outside the repository, with your secrets manager (External Secrets, Sealed Secrets…).
- Monitoring: Keycloak exposes metrics (
metrics-enabled) and the operator creates aServiceMonitorwhen the Prometheus Operator is installed. CloudNativePG exposes its own. - Upgrades: upgrading the operator upgrades Keycloak. Read the migration notes, back up the database, test on a non-production cluster first.
In 2022, deploying Keycloak on Kubernetes meant writing a Deployment by hand and hoping the cluster formed. In 2026, three operators do the heavy lifting: CloudNativePG for a replicated PostgreSQL that fails over by itself, the Keycloak Operator for clustered, rolling-updated Keycloak pods, and Envoy Gateway for a single entry point. With the Gateway API comes a bonus: Envoy Gateway also protects applications that were never designed for OpenID Connect, without changing a line of their code.
None of it makes Kubernetes the right choice for every Keycloak, though. Choose it when your platform already lives there, and otherwise, let someone else run Keycloak for you.
Whichever way you go, we can help you design it: Keycloak architecture and Keycloak support.
Repository: https://github.com/please-openit/keycloak-kubernetes