Decision. CloudNativePG operator
Alternative. A hand-rolled StatefulSet
Why. The operator owns failover, replica promotion, and Barman backups natively; hand-rolling that is a second job.
PostgreSQL as a primary with replicas, backed up to object storage via the CNPG operator
The database layer is one umbrella Application: the CloudNativePG operator, pgAdmin, and PgBouncer land together. Clusters run primary plus replicas, and backups go to object storage through the Barman integration built into the operator. Velero deliberately skips these namespaces.
Excerpt from argocd-apps/values.yaml in the argocd-apps chart,
with annotations added for this site:
postgres:
application: true
project: cluster-services
# one umbrella Application: operator + pgAdmin + PgBouncer land together
info: "This application includes cloudnative-pg operator, pg-admin and pg-bouncer"
# false
autoSync: false
The full postgres/values.yaml from the service's
own chart:
# CloudNativePG operator subchart values.
cloudnative-pg:
replicaCount: 1
image:
repository: ghcr.io/cloudnative-pg/cloudnative-pg
pullPolicy: IfNotPresent
tag: 1.30.0
crds:
create: true
webhook:
port: 9443
mutating:
create: true
failurePolicy: Fail
validating:
create: true
failurePolicy: Fail
livenessProbe:
initialDelaySeconds: 3
readinessProbe:
initialDelaySeconds: 3
config:
create: true
serviceAccount:
create: true
name: "postgres-sa"
rbac:
create: true
aggregateClusterRoles: false
containerSecurityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
runAsUser: 10001
runAsGroup: 10001
seccompProfile:
type: RuntimeDefault
capabilities:
drop:
- "ALL"
podSecurityContext:
runAsNonRoot: true
seccompProfile:
type: RuntimeDefault
service:
type: ClusterIP
name: cnpg-webhook-service # DO NOT CHANGE
port: 443
# Barman Cloud plugin subchart values.
plugin-barman-cloud:
replicaCount: 1
image:
registry: ghcr.io
repository: cloudnative-pg/plugin-barman-cloud
pullPolicy: IfNotPresent
sidecarImage:
registry: ghcr.io
repository: cloudnative-pg/plugin-barman-cloud-sidecar
crds:
create: true
# PostgreSQL cluster definition.
cluster:
# Number of instances (primary + replicas).
instances: 3
# PostgreSQL 18 (bare major tag, per CNPG container image convention).
imageName: ghcr.io/cloudnative-pg/postgresql:18
imagePullPolicy: IfNotPresent
postgresUID: 26
postgresGID: 26
# Rolling update behaviour: switch over the primary after replicas are updated.
primaryUpdateMethod: switchover
primaryUpdateStrategy: unsupervised
logLevel: "info"
stopDelay: 1800
smartShutdownTimeout: 180
enablePDB: true
enableSuperuserAccess: true
superuserSecret: postgres-superuser-secret
# CNPG AffinityConfiguration. The MINISFORUM is a *preference*, not a hard
# pin: nodeAffinity.preferred lets the primary land there while still
# allowing CNPG to schedule/fail over onto other nodes if it is unavailable.
affinity:
topologyKey: kubernetes.io/hostname
nodeAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
preference:
matchExpressions:
- key: kubernetes.io/hostname
operator: In
values:
- opi5-worker-5-genai
# Soft spread of the 3 instances across distinct hostnames. ScheduleAnyway
# keeps the replicas schedulable even if the MINISFORUM is the only matching
# node, so failover can still escape to another node.
# The template injects the `cnpg.io/cluster` labelSelector automatically.
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
storage:
size: 50Gi
storageClass: local-path-ssd
# WAL lives on the same PVC by default (colocated with PGDATA). No separate
# walStorage unless performance demands it.
postgresql:
parameters:
shared_buffers: 256MB
pg_stat_statements.max: "10000"
pg_stat_statements.track: all
auto_explain.log_min_duration: "10s"
pg_hba:
- "local all all scram-sha-256"
# Pod CIDR (10.42.0.0/16) — applications and poolers connect from pod IPs.
# Service CIDR (10.43.0.0/16) included for completeness.
- "host all all 10.42.0.0/16 scram-sha-256"
- "host all all 10.43.0.0/16 scram-sha-256"
- "host all all 0.0.0.0/0 reject"
bootstrap:
initdb:
database: opi5-cluster
owner: opi5-cluster
secret:
name: postgres-user-secret
# Barman Cloud plugin for WAL archiving + base backups.
# barmanObjectName references the ObjectStore resource rendered by
# templates/cluster/objectstore.yaml.
plugins:
- name: barman-cloud.cloudnative-pg.io
isWALArchiver: true
parameters:
barmanObjectName: postgres-backup-store
# Declarative databases managed via the Database CRD.
# `name` is the PostgreSQL database identifier (underscores); the Kubernetes
# resource name is derived by replacing `_` with `-`.
databases:
- name: argo_workflows
- name: metabase
- name: open_webui
- name: payload_cms
- name: bazarr
- name: prowlarr_logs
- name: prowlarr
- name: radarr_logs
- name: radarr
- name: sonarr_logs
- name: sonarr
- name: bouncer
- name: beacon
- name: quote_my_shizzle
- name: meridian
- name: muse
- name: whistle
- name: wos_assistant
- name: trove
extensions:
- name: vector
# Backup and disaster recovery via the Barman Cloud plugin.
# Destination: standalone RustFS on backup-raspi3 (LAN, S3-compatible,
# plain HTTP — accepted posture, same as the velero effort).
# The postgres-backups bucket on that server is a NEW bucket (the old R2
# bucket of the same name is retired after a 30d wind-down).
backups:
enabled: true
endpointURL: http://backup-raspi3.opi5cluster.co.uk:9000
destinationPath: s3://postgres-backups/
# RustFS on backup-raspi3. Keys come from Vault key `rustfs`
# (properties backup-access-key / backup-secret-key) via ESO — see
# templates/secrets/rustfs-backup-external-secret.yaml, which renders
# the rustfs-backup-secret (keys: access-key / secret-key).
s3CredentialsSecret: rustfs-backup-secret
s3CredentialsSecretKeys:
accessKeyId: access-key
secretAccessKey: secret-key
retentionPolicy: "30d"
wal:
compression: gzip
data:
compression: gzip
additionalCommandArgs:
- "--min-chunk-size=5MB"
- "--read-timeout=60"
scheduledBackups:
- name: daily-backup
schedule: "0 0 0 * * *"
backupOwnerReference: self
# PgBouncer connection pooler(s).
poolers:
- name: rw
type: rw
poolMode: session
instances: 1
parameters:
max_client_conn: "1000"
default_pool_size: "50"
template:
metadata:
labels:
app: pooler
spec:
containers:
- name: pgbouncer
image: ghcr.io/cloudnative-pg/pgbouncer:1.25.2
# Bundled Rancher Local Path Provisioner lives in the longhorn project.
# The `local-path-ssd` StorageClass it provisions is consumed here via
# cluster.storage.storageClass and pgadmin.persistentVolumeClaim.storageClassName.
pgadmin:
name: pgadmin
image: dpage/pgadmin4:9.17
userEmail: tech@opi5cluster.co.uk
port: 80
securityContext:
runAsUser: 5050
runAsGroup: 5050
persistentVolumeClaim:
accessModes: ReadWriteOnce
storage: 3Gi
storageClassName: local-path-ssd
resources:
requests:
cpu: 100m
memory: 128M
limits:
cpu: 300m
memory: 547M Decision. CloudNativePG operator
Alternative. A hand-rolled StatefulSet
Why. The operator owns failover, replica promotion, and Barman backups natively; hand-rolling that is a second job.
Decision. Exclude postgres namespaces from Velero
Alternative. Letting Velero and the operator both back up the database
Why. Two systems racing on the same database invites inconsistent restores; the CNPG path is the authoritative one.