Velero

Whole-cluster backups: nightly manifests and PVC data (Kopia) to RustFS running on backup-raspi3, kept 30 days

Velero owns disaster recovery for the whole cluster. Two schedules run nightly: manifests only at 02:00, and file-level application data at 03:00 across every namespace except postgres and velero itself. Data is shipped with the Kopia file-level action to the velero bucket on a standalone RustFS instance running on backup-raspi3, a device outside the cluster, with a 30-day TTL.

ArgoCD configuration

Excerpt from argocd-apps/values.yaml in the argocd-apps chart, with annotations added for this site:

velero:
  # true
  application: true
  project: k3s-services
  # false: a bad Velero upgrade should never surface mid-backup
  autoSync: false

Chart values

The full velero/values.yaml from the service's own chart:

velero:
  image:
    tag: v1.18.1 # match the vendored chart appVersion

  initContainers:
    - name: velero-plugin-for-aws
      image: velero/velero-plugin-for-aws:v1.13.1
      imagePullPolicy: IfNotPresent
      volumeMounts:
        - mountPath: /target
          name: plugins

  # Credentials are NOT rendered by the chart — the BSL references the
  # ESO-managed velero-cloud-credentials secret directly (templates/external-secret.yaml).
  credentials:
    useSecret: false

  configuration:
    backupStorageLocation:
      - name: default
        provider: aws
        bucket: velero
        default: true
        credential:
          name: velero-cloud-credentials
          key: cloud
        config:
          region: minio
          s3ForcePathStyle: "true"
          s3Url: http://backup-raspi3.opi5cluster.co.uk:9000
    defaultVolumesToFsBackup: false # only the daily-app-data schedule opts in

  deployNodeAgent: true

  nodeAgent:
    tolerations:
      - key: opi5.cluster/role # data/genai tainted nodes host PVCs
        operator: Equal
        value: data
        effect: NoSchedule
      - key: node-role.kubernetes.io/master # control plane: node-agent must run there to back up pod volumes (metallb-frr et al)
        operator: Exists
        effect: NoSchedule
      - key: node.kubernetes.io/not-ready
        operator: Exists
        effect: NoSchedule
      - key: node.kubernetes.io/unreachable
        operator: Exists
        effect: NoSchedule
    resources:
      requests:
        cpu: 100m
        memory: 128Mi

  resources:
    requests:
      cpu: 100m
      memory: 128Mi


  # No CSI snapshot locations; Kopia file-level backup handles volume data.
  snapshotsEnabled: false

  schedules:
    # All namespaces, metadata only — fast DR catalog of the whole cluster.
    daily-manifests:
      schedule: "0 2 * * *"
      template:
        ttl: 720h
        defaultVolumesToFsBackup: false
        snapshotVolumes: false
    # All namespaces except postgres (CNPG/Barman owns that pipeline) and
    # velero itself, with Kopia file-level backup of every PVC.
    daily-app-data:
      schedule: "0 3 * * *"
      template:
        ttl: 720h
        defaultVolumesToFsBackup: true
        snapshotVolumes: false
        excludedNamespaces:
          - postgres
          - velero

# -----------------------------------------------------------------------------
# velero-ui (otwld/velero-ui 0.15.0, app 0.10.2) — authenticated web dashboard
# over this Velero instance. Served at velero.opi5cluster.co.uk via our HTTPRoute;
# the Velero server itself stays headless (see AGENTS.md rule 6).
velero-ui:
  image:
    tag: "0.10.2" # pin to the chart's appVersion

  # The JWT signing passphrase via ESO secret (env below). useSecret must
  # stay false: the chart otherwise renders a Secret holding its well-known
  # default passphrase, and never wires it into the pod anyway.
  configuration:
    general:
      secretPassPhrase:
        useSecret: false

  rbac:
    create: true
    clusterAdministrator: false # NEVER bind cluster-admin; keeps the chart's velero.io ClusterRole + namespaced Roles instead

  env:
    - name: BASIC_AUTH_USERNAME
      value: admin
    - name: BASIC_AUTH_PASSWORD
      valueFrom:
        secretKeyRef:
          name: velero-ui-credentials
          key: admin-password
    - name: AUTH_SECRET_PASSPHRASE
      valueFrom:
        secretKeyRef:
          name: velero-ui-credentials
          key: auth-passphrase

  resources:
    requests:
      cpu: 50m
      memory: 128Mi
    limits:
      cpu: 500m
      memory: 512Mi

  # Consumed by templates/velero-ui-httproute.yaml (our template — the
  # subchart ignores unknown keys). Requires the AdGuard LAN DNS record.
  hostname: velero.opi5cluster.co.uk

Manifests & templates

templates/external-secret.yaml

ExternalSecret syncing the service credentials from Vault into the namespace.

Show manifest
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
  name: velero-cloud-credentials
  namespace: {{ .Release.Namespace }}
spec:
  refreshInterval: 1h
  secretStoreRef:
    kind: ClusterSecretStore
    name: vault-cluster-secret-store
  target:
    name: velero-cloud-credentials
    creationPolicy: Owner
    template:
      data:
        # AWS ini consumed by the velero aws plugin. The inner brace pairs
        # below are ESO template expressions — escaped from Helm via raw
        # literals and resolved by ESO from the data map below. Hyphenated
        # keys require `index`; dotted access can't parse them (and Helm
        # would eat the braces first, yielding empty creds → BSL falls
        # back to IMDS).
        cloud: |
          [default]
          aws_access_key_id={{ `{{ index . "backup-access-key" }}` }}
          aws_secret_access_key={{ `{{ index . "backup-secret-key" }}` }}
  data:
    - secretKey: backup-access-key
      remoteRef:
        key: rustfs
        property: backup-access-key
    - secretKey: backup-secret-key
      remoteRef:
        key: rustfs
        property: backup-secret-key

templates/velero-ui-externalsecret.yaml

ExternalSecret for the Velero UI login.

Show manifest
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
  name: velero-ui-credentials
  namespace: {{ .Release.Namespace }}
spec:
  refreshInterval: 1h
  secretStoreRef:
    kind: ClusterSecretStore
    name: vault-cluster-secret-store
  target:
    name: velero-ui-credentials
    creationPolicy: Owner
  data:
    # UI login (BASIC_AUTH_USERNAME=admin) and the JWT session-signing
    # passphrase — two distinct secrets, two distinct rotation blast radii.
    - secretKey: admin-password
      remoteRef:
        key: velero
        property: admin-password
    - secretKey: auth-passphrase
      remoteRef:
        key: velero
        property: auth-passphrase

templates/velero-ui-httproute.yaml

HTTPRoute for the Velero UI.

Show manifest
{{- $vui := index .Values "velero-ui" }}
# TLS terminates on the gateway's `websecure` listener — routes carry no TLS
# block of their own (see homepage/longhorn routes).
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: velero-ui-http
  namespace: {{ .Release.Namespace }}
  annotations:
    link.argocd.argoproj.io/external-link: https://{{ $vui.hostname }}
spec:
  parentRefs:
    - name: istio-gateway
      namespace: istio
      sectionName: websecure
  hostnames:
    - {{ $vui.hostname | quote }}
  rules:
    - matches:
        - path:
            type: PathPrefix
            value: /
      backendRefs:
        # Mirror of the velero-ui chart's fullname helper for release `velero`:
        # `velero-velero-ui`; its Service listens on service.port (3000).
        - name: {{ printf "%s-%s" .Release.Name "velero-ui" }}
          port: 3000
  • Postgres is excluded because CloudNativePG runs its own backup lifecycle (Barman to object storage).

Trade-offs

Decision. One whole-cluster scope with two schedules

Alternative. Per-application backup configuration

Why. One restore path for disaster recovery beats per-app logic that rots; exclusions are explicit and few.

Decision. File-level Kopia backup instead of CSI volume snapshots

Alternative. CSI snapshots on Longhorn

Why. Snapshots alone would live on the same storage as the data; the Kopia path ships bytes off-cluster to a different device.

Decision. Back up to a device outside the cluster

Alternative. Backing up to in-cluster object storage

Why. A backup that dies with the cluster is not a backup.

← Back to Platform & Infrastructure · All service groups