Decision. Longhorn
Alternative. Rook/Ceph
Why. Rejected early (see the Lessons page): Ceph operational overhead and arm64 quirks are not worth it at this scale, while Longhorn is a UI away from understandable.
Replicates block storage across nodes, so a lost disk is not a lost volume
Longhorn is the default storage class: every PVC is replicated across nodes, so any single node can disappear without losing data. Volumes are created on demand, snapshotted on a schedule, and exposed over iSCSI under the hood.
Excerpt from argocd-apps/values.yaml in the argocd-apps chart,
with annotations added for this site:
longhorn:
# true
application: true
project: k3s-services
# false: storage engine upgrades are deliberate events
autoSync: false
The full longhorn/values.yaml from the service's
own chart:
longhorn:
service:
ui:
type: ClusterIP
manager:
type: ClusterIP
#defaultSettings:
# backupTarget: s3://opi5-cluster-longhorn@eu-west-2/backups
# backupTargetCredentialSecret: longhorn-backup-secret
# defaultDataPath: /mnt/ssd-small/storage
# System-managed components (engine-image, instance-manager, CSI) need
# the same data-role toleration as longhornManager, or the engine binary
# is never deployed to tainted data/genai nodes and the backup-target
# controller logs "failed to execute ... no such file or directory"
# every 5 minutes.
defaultSettings:
taintToleration: opi5.cluster/role=data:NoSchedule
persistence:
defaultClass: false
longhornManager:
log:
format: json
tolerations:
- key: "opi5.cluster/role"
operator: "Equal"
value: "data"
effect: "NoSchedule"
ingress:
enabled: false
# Bundled Rancher Local Path Provisioner.
# k3s' default `local-path` provisioner is not installed, so this chart ships
# its own single Deployment (with helper pods) writing PVCs to
# /mnt/ssd-small/local on every node.
#
# Consumed by the postgres project via the `local-path-ssd` StorageClass for
# PostgreSQL PGDATA (no storage-level replication — CNPG handles it).
localPathProvisioner:
enabled: true
# v0.0.37+ required: it is the first release serving the /health and /ready
# endpoints used by the startup/liveness/readiness probes in the
# local-path-provisioner Deployment. Older images (e.g. v0.0.30) have no
# health server, so the probes fail with "connection refused".
image: rancher/local-path-provisioner:v0.0.37
nodePath: /mnt/ssd-small/local
# Tolerations for the per-node helper pods that create/remove the volume
# directories. Must include the cluster's data/genai node taint so the
# helper pod can run on the node hosting the PVC.
helperPodTolerations:
- key: node.kubernetes.io/not-ready
operator: Exists
effect: NoSchedule
- key: node.kubernetes.io/disk-pressure
operator: Exists
effect: NoSchedule
- key: opi5.cluster/role
operator: Equal
value: data
effect: NoSchedule
# Tolerations for the provisioner Deployment pod itself.
tolerations:
- key: opi5.cluster/role
operator: Equal
value: data
effect: NoSchedule
storageClass:
name: local-path-ssd
reclaimPolicy: Retain
defaultClass: false templates/http-route.yaml Gateway API HTTPRoute exposing the service through the Istio gateway under opi5cluster.co.uk.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: longhorn-httproute
namespace: longhorn
annotations:
link.argocd.argoproj.io/external-link: "https://longhorn.opi5cluster.co.uk"
spec:
parentRefs:
- name: istio-gateway
namespace: istio
sectionName: websecure
hostnames:
- longhorn.opi5cluster.co.uk
rules:
- backendRefs:
- name: longhorn-frontend
port: 80 templates/storage-class-ssd-large.yaml Longhorn StorageClass (ssd-large): replica count and replica placement tuning.
kind: StorageClass
apiVersion: storage.k8s.io/v1
metadata:
name: longhorn-ssd-large
annotations:
longhorn.io/last-applied-configmap: |
kind: StorageClass
apiVersion: storage.k8s.io/v1
metadata:
name: longhorn-ssd-large
annotations:
storageclass.kubernetes.io/is-default-class: "false"
storageclass.kubernetes.io/is-default-class: 'false'
provisioner: driver.longhorn.io
allowVolumeExpansion: true
reclaimPolicy: Retain
volumeBindingMode: Immediate
parameters:
numberOfReplicas: "2"
staleReplicaTimeout: "2880"
fsType: "xfs"
dataLocality: "best-effort"
diskSelector: "ssd-large" templates/storage-class-ssd-small.yaml Longhorn StorageClass (ssd-small): replica count and replica placement tuning.
kind: StorageClass
apiVersion: storage.k8s.io/v1
metadata:
name: longhorn-ssd-small
annotations:
longhorn.io/last-applied-configmap: |
kind: StorageClass
apiVersion: storage.k8s.io/v1
metadata:
name: longhorn-ssd-small
annotations:
storageclass.kubernetes.io/is-default-class: "true"
storageclass.kubernetes.io/is-default-class: 'true'
provisioner: driver.longhorn.io
allowVolumeExpansion: true
reclaimPolicy: Retain
volumeBindingMode: Immediate
parameters:
numberOfReplicas: "2"
staleReplicaTimeout: "2880"
fsType: "ext4"
dataLocality: "best-effort"
diskSelector: "ssd-small" Decision. Longhorn
Alternative. Rook/Ceph
Why. Rejected early (see the Lessons page): Ceph operational overhead and arm64 quirks are not worth it at this scale, while Longhorn is a UI away from understandable.