Decision. One whole-cluster scope with two schedules
Alternative. Per-application backup configuration
Why. One restore path for disaster recovery beats per-app logic that rots; exclusions are explicit and few.
Whole-cluster backups: nightly manifests and PVC data (Kopia) to RustFS running on backup-raspi3, kept 30 days
Velero owns disaster recovery for the whole cluster. Two schedules run nightly: manifests only at 02:00, and file-level application data at 03:00 across every namespace except postgres and velero itself. Data is shipped with the Kopia file-level action to the velero bucket on a standalone RustFS instance running on backup-raspi3, a device outside the cluster, with a 30-day TTL.
Excerpt from argocd-apps/values.yaml in the argocd-apps chart,
with annotations added for this site:
velero:
# true
application: true
project: k3s-services
# false: a bad Velero upgrade should never surface mid-backup
autoSync: false
The full velero/values.yaml from the service's
own chart:
velero:
image:
tag: v1.18.1 # match the vendored chart appVersion
initContainers:
- name: velero-plugin-for-aws
image: velero/velero-plugin-for-aws:v1.13.1
imagePullPolicy: IfNotPresent
volumeMounts:
- mountPath: /target
name: plugins
# Credentials are NOT rendered by the chart — the BSL references the
# ESO-managed velero-cloud-credentials secret directly (templates/external-secret.yaml).
credentials:
useSecret: false
configuration:
backupStorageLocation:
- name: default
provider: aws
bucket: velero
default: true
credential:
name: velero-cloud-credentials
key: cloud
config:
region: minio
s3ForcePathStyle: "true"
s3Url: http://backup-raspi3.opi5cluster.co.uk:9000
defaultVolumesToFsBackup: false # only the daily-app-data schedule opts in
deployNodeAgent: true
nodeAgent:
tolerations:
- key: opi5.cluster/role # data/genai tainted nodes host PVCs
operator: Equal
value: data
effect: NoSchedule
- key: node-role.kubernetes.io/master # control plane: node-agent must run there to back up pod volumes (metallb-frr et al)
operator: Exists
effect: NoSchedule
- key: node.kubernetes.io/not-ready
operator: Exists
effect: NoSchedule
- key: node.kubernetes.io/unreachable
operator: Exists
effect: NoSchedule
resources:
requests:
cpu: 100m
memory: 128Mi
resources:
requests:
cpu: 100m
memory: 128Mi
# No CSI snapshot locations; Kopia file-level backup handles volume data.
snapshotsEnabled: false
schedules:
# All namespaces, metadata only — fast DR catalog of the whole cluster.
daily-manifests:
schedule: "0 2 * * *"
template:
ttl: 720h
defaultVolumesToFsBackup: false
snapshotVolumes: false
# All namespaces except postgres (CNPG/Barman owns that pipeline) and
# velero itself, with Kopia file-level backup of every PVC.
daily-app-data:
schedule: "0 3 * * *"
template:
ttl: 720h
defaultVolumesToFsBackup: true
snapshotVolumes: false
excludedNamespaces:
- postgres
- velero
# -----------------------------------------------------------------------------
# velero-ui (otwld/velero-ui 0.15.0, app 0.10.2) — authenticated web dashboard
# over this Velero instance. Served at velero.opi5cluster.co.uk via our HTTPRoute;
# the Velero server itself stays headless (see AGENTS.md rule 6).
velero-ui:
image:
tag: "0.10.2" # pin to the chart's appVersion
# The JWT signing passphrase via ESO secret (env below). useSecret must
# stay false: the chart otherwise renders a Secret holding its well-known
# default passphrase, and never wires it into the pod anyway.
configuration:
general:
secretPassPhrase:
useSecret: false
rbac:
create: true
clusterAdministrator: false # NEVER bind cluster-admin; keeps the chart's velero.io ClusterRole + namespaced Roles instead
env:
- name: BASIC_AUTH_USERNAME
value: admin
- name: BASIC_AUTH_PASSWORD
valueFrom:
secretKeyRef:
name: velero-ui-credentials
key: admin-password
- name: AUTH_SECRET_PASSPHRASE
valueFrom:
secretKeyRef:
name: velero-ui-credentials
key: auth-passphrase
resources:
requests:
cpu: 50m
memory: 128Mi
limits:
cpu: 500m
memory: 512Mi
# Consumed by templates/velero-ui-httproute.yaml (our template — the
# subchart ignores unknown keys). Requires the AdGuard LAN DNS record.
hostname: velero.opi5cluster.co.uk templates/external-secret.yaml ExternalSecret syncing the service credentials from Vault into the namespace.
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
name: velero-cloud-credentials
namespace: {{ .Release.Namespace }}
spec:
refreshInterval: 1h
secretStoreRef:
kind: ClusterSecretStore
name: vault-cluster-secret-store
target:
name: velero-cloud-credentials
creationPolicy: Owner
template:
data:
# AWS ini consumed by the velero aws plugin. The inner brace pairs
# below are ESO template expressions — escaped from Helm via raw
# literals and resolved by ESO from the data map below. Hyphenated
# keys require `index`; dotted access can't parse them (and Helm
# would eat the braces first, yielding empty creds → BSL falls
# back to IMDS).
cloud: |
[default]
aws_access_key_id={{ `{{ index . "backup-access-key" }}` }}
aws_secret_access_key={{ `{{ index . "backup-secret-key" }}` }}
data:
- secretKey: backup-access-key
remoteRef:
key: rustfs
property: backup-access-key
- secretKey: backup-secret-key
remoteRef:
key: rustfs
property: backup-secret-key templates/velero-ui-externalsecret.yaml ExternalSecret for the Velero UI login.
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
name: velero-ui-credentials
namespace: {{ .Release.Namespace }}
spec:
refreshInterval: 1h
secretStoreRef:
kind: ClusterSecretStore
name: vault-cluster-secret-store
target:
name: velero-ui-credentials
creationPolicy: Owner
data:
# UI login (BASIC_AUTH_USERNAME=admin) and the JWT session-signing
# passphrase — two distinct secrets, two distinct rotation blast radii.
- secretKey: admin-password
remoteRef:
key: velero
property: admin-password
- secretKey: auth-passphrase
remoteRef:
key: velero
property: auth-passphrase templates/velero-ui-httproute.yaml HTTPRoute for the Velero UI.
{{- $vui := index .Values "velero-ui" }}
# TLS terminates on the gateway's `websecure` listener — routes carry no TLS
# block of their own (see homepage/longhorn routes).
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: velero-ui-http
namespace: {{ .Release.Namespace }}
annotations:
link.argocd.argoproj.io/external-link: https://{{ $vui.hostname }}
spec:
parentRefs:
- name: istio-gateway
namespace: istio
sectionName: websecure
hostnames:
- {{ $vui.hostname | quote }}
rules:
- matches:
- path:
type: PathPrefix
value: /
backendRefs:
# Mirror of the velero-ui chart's fullname helper for release `velero`:
# `velero-velero-ui`; its Service listens on service.port (3000).
- name: {{ printf "%s-%s" .Release.Name "velero-ui" }}
port: 3000 Decision. One whole-cluster scope with two schedules
Alternative. Per-application backup configuration
Why. One restore path for disaster recovery beats per-app logic that rots; exclusions are explicit and few.
Decision. File-level Kopia backup instead of CSI volume snapshots
Alternative. CSI snapshots on Longhorn
Why. Snapshots alone would live on the same storage as the data; the Kopia path ships bytes off-cluster to a different device.
Decision. Back up to a device outside the cluster
Alternative. Backing up to in-cluster object storage
Why. A backup that dies with the cluster is not a backup.