PostgreSQL (CloudNativePG)

PostgreSQL as a primary with replicas, backed up to object storage via the CNPG operator

The database layer is one umbrella Application: the CloudNativePG operator, pgAdmin, and PgBouncer land together. Clusters run primary plus replicas, and backups go to object storage through the Barman integration built into the operator. Velero deliberately skips these namespaces.

ArgoCD configuration

Excerpt from argocd-apps/values.yaml in the argocd-apps chart, with annotations added for this site:

postgres:
  application: true
  project: cluster-services
  # one umbrella Application: operator + pgAdmin + PgBouncer land together
  info: "This application includes cloudnative-pg operator, pg-admin and pg-bouncer"
  # false
  autoSync: false

Chart values

The full postgres/values.yaml from the service's own chart:

# CloudNativePG operator subchart values.
cloudnative-pg:
  replicaCount: 1
  image:
    repository: ghcr.io/cloudnative-pg/cloudnative-pg
    pullPolicy: IfNotPresent
    tag: 1.30.0

  crds:
    create: true

  webhook:
    port: 9443
    mutating:
      create: true
      failurePolicy: Fail
    validating:
      create: true
      failurePolicy: Fail
    livenessProbe:
      initialDelaySeconds: 3
    readinessProbe:
      initialDelaySeconds: 3

  config:
    create: true

  serviceAccount:
    create: true
    name: "postgres-sa"

  rbac:
    create: true
    aggregateClusterRoles: false

  containerSecurityContext:
    allowPrivilegeEscalation: false
    readOnlyRootFilesystem: true
    runAsUser: 10001
    runAsGroup: 10001
    seccompProfile:
      type: RuntimeDefault
    capabilities:
      drop:
        - "ALL"

  podSecurityContext:
    runAsNonRoot: true
    seccompProfile:
      type: RuntimeDefault

  service:
    type: ClusterIP
    name: cnpg-webhook-service # DO NOT CHANGE
    port: 443

# Barman Cloud plugin subchart values.
plugin-barman-cloud:
  replicaCount: 1
  image:
    registry: ghcr.io
    repository: cloudnative-pg/plugin-barman-cloud
    pullPolicy: IfNotPresent
  sidecarImage:
    registry: ghcr.io
    repository: cloudnative-pg/plugin-barman-cloud-sidecar
  crds:
    create: true

# PostgreSQL cluster definition.
cluster:
  # Number of instances (primary + replicas).
  instances: 3
  # PostgreSQL 18 (bare major tag, per CNPG container image convention).
  imageName: ghcr.io/cloudnative-pg/postgresql:18
  imagePullPolicy: IfNotPresent

  postgresUID: 26
  postgresGID: 26

  # Rolling update behaviour: switch over the primary after replicas are updated.
  primaryUpdateMethod: switchover
  primaryUpdateStrategy: unsupervised

  logLevel: "info"
  stopDelay: 1800
  smartShutdownTimeout: 180
  enablePDB: true

  enableSuperuserAccess: true
  superuserSecret: postgres-superuser-secret

  # CNPG AffinityConfiguration. The MINISFORUM is a *preference*, not a hard
  # pin: nodeAffinity.preferred lets the primary land there while still
  # allowing CNPG to schedule/fail over onto other nodes if it is unavailable.
  affinity:
    topologyKey: kubernetes.io/hostname
    nodeAffinity:
      preferredDuringSchedulingIgnoredDuringExecution:
        - weight: 100
          preference:
            matchExpressions:
              - key: kubernetes.io/hostname
                operator: In
                values:
                  - opi5-worker-5-genai

  # Soft spread of the 3 instances across distinct hostnames. ScheduleAnyway
  # keeps the replicas schedulable even if the MINISFORUM is the only matching
  # node, so failover can still escape to another node.
  # The template injects the `cnpg.io/cluster` labelSelector automatically.
  topologySpreadConstraints:
    - maxSkew: 1
      topologyKey: kubernetes.io/hostname
      whenUnsatisfiable: ScheduleAnyway

  storage:
    size: 50Gi
    storageClass: local-path-ssd

  # WAL lives on the same PVC by default (colocated with PGDATA). No separate
  # walStorage unless performance demands it.

  postgresql:
    parameters:
      shared_buffers: 256MB
      pg_stat_statements.max: "10000"
      pg_stat_statements.track: all
      auto_explain.log_min_duration: "10s"
    pg_hba:
      - "local all all scram-sha-256"
      # Pod CIDR (10.42.0.0/16) — applications and poolers connect from pod IPs.
      # Service CIDR (10.43.0.0/16) included for completeness.
      - "host  all all 10.42.0.0/16    scram-sha-256"
      - "host  all all 10.43.0.0/16    scram-sha-256"
      - "host  all all 0.0.0.0/0       reject"

  bootstrap:
    initdb:
      database: opi5-cluster
      owner: opi5-cluster
      secret:
        name: postgres-user-secret

  # Barman Cloud plugin for WAL archiving + base backups.
  # barmanObjectName references the ObjectStore resource rendered by
  # templates/cluster/objectstore.yaml.
  plugins:
    - name: barman-cloud.cloudnative-pg.io
      isWALArchiver: true
      parameters:
        barmanObjectName: postgres-backup-store

# Declarative databases managed via the Database CRD.
# `name` is the PostgreSQL database identifier (underscores); the Kubernetes
# resource name is derived by replacing `_` with `-`.
databases:
  - name: argo_workflows
  - name: metabase
  - name: open_webui
  - name: payload_cms
  - name: bazarr
  - name: prowlarr_logs
  - name: prowlarr
  - name: radarr_logs
  - name: radarr
  - name: sonarr_logs
  - name: sonarr
  - name: bouncer
  - name: beacon
  - name: quote_my_shizzle
  - name: meridian
  - name: muse
  - name: whistle
  - name: wos_assistant
  - name: trove
    extensions:
      - name: vector

# Backup and disaster recovery via the Barman Cloud plugin.
# Destination: standalone RustFS on backup-raspi3 (LAN, S3-compatible,
# plain HTTP — accepted posture, same as the velero effort).
# The postgres-backups bucket on that server is a NEW bucket (the old R2
# bucket of the same name is retired after a 30d wind-down).
backups:
  enabled: true
  endpointURL: http://backup-raspi3.opi5cluster.co.uk:9000
  destinationPath: s3://postgres-backups/
  # RustFS on backup-raspi3. Keys come from Vault key `rustfs`
  # (properties backup-access-key / backup-secret-key) via ESO — see
  # templates/secrets/rustfs-backup-external-secret.yaml, which renders
  # the rustfs-backup-secret (keys: access-key / secret-key).
  s3CredentialsSecret: rustfs-backup-secret
  s3CredentialsSecretKeys:
    accessKeyId: access-key
    secretAccessKey: secret-key
  retentionPolicy: "30d"
  wal:
    compression: gzip
  data:
    compression: gzip
    additionalCommandArgs:
      - "--min-chunk-size=5MB"
      - "--read-timeout=60"
  scheduledBackups:
    - name: daily-backup
      schedule: "0 0 0 * * *"
      backupOwnerReference: self

# PgBouncer connection pooler(s).
poolers:
  - name: rw
    type: rw
    poolMode: session
    instances: 1
    parameters:
      max_client_conn: "1000"
      default_pool_size: "50"
    template:
      metadata:
        labels:
          app: pooler
      spec:
        containers:
          - name: pgbouncer
            image: ghcr.io/cloudnative-pg/pgbouncer:1.25.2

# Bundled Rancher Local Path Provisioner lives in the longhorn project.
# The `local-path-ssd` StorageClass it provisions is consumed here via
# cluster.storage.storageClass and pgadmin.persistentVolumeClaim.storageClassName.

pgadmin:
  name: pgadmin
  image: dpage/pgadmin4:9.17
  userEmail: tech@opi5cluster.co.uk
  port: 80
  securityContext:
    runAsUser: 5050
    runAsGroup: 5050
  persistentVolumeClaim:
    accessModes: ReadWriteOnce
    storage: 3Gi
    storageClassName: local-path-ssd
  resources:
    requests:
      cpu: 100m
      memory: 128M
    limits:
      cpu: 300m
      memory: 547M

Trade-offs

Decision. CloudNativePG operator

Alternative. A hand-rolled StatefulSet

Why. The operator owns failover, replica promotion, and Barman backups natively; hand-rolling that is a second job.

Decision. Exclude postgres namespaces from Velero

Alternative. Letting Velero and the operator both back up the database

Why. Two systems racing on the same database invites inconsistent restores; the CNPG path is the authoritative one.

← Back to Data & Streaming · All service groups