Longhorn

Replicates block storage across nodes, so a lost disk is not a lost volume

Longhorn is the default storage class: every PVC is replicated across nodes, so any single node can disappear without losing data. Volumes are created on demand, snapshotted on a schedule, and exposed over iSCSI under the hood.

ArgoCD configuration

Excerpt from argocd-apps/values.yaml in the argocd-apps chart, with annotations added for this site:

longhorn:
  # true
  application: true
  project: k3s-services
  # false: storage engine upgrades are deliberate events
  autoSync: false

Chart values

The full longhorn/values.yaml from the service's own chart:

longhorn:
  service:
    ui:
      type: ClusterIP
    manager:
      type: ClusterIP

  #defaultSettings:
  #  backupTarget: s3://opi5-cluster-longhorn@eu-west-2/backups
  #  backupTargetCredentialSecret: longhorn-backup-secret
  #  defaultDataPath: /mnt/ssd-small/storage

  # System-managed components (engine-image, instance-manager, CSI) need
  # the same data-role toleration as longhornManager, or the engine binary
  # is never deployed to tainted data/genai nodes and the backup-target
  # controller logs "failed to execute ... no such file or directory"
  # every 5 minutes.
  defaultSettings:
    taintToleration: opi5.cluster/role=data:NoSchedule

  persistence:
    defaultClass: false

  longhornManager:
    log:
      format: json

    tolerations:
      - key: "opi5.cluster/role"
        operator: "Equal"
        value: "data"
        effect: "NoSchedule"

  ingress:
    enabled: false

# Bundled Rancher Local Path Provisioner.
# k3s' default `local-path` provisioner is not installed, so this chart ships
# its own single Deployment (with helper pods) writing PVCs to
# /mnt/ssd-small/local on every node.
#
# Consumed by the postgres project via the `local-path-ssd` StorageClass for
# PostgreSQL PGDATA (no storage-level replication — CNPG handles it).
localPathProvisioner:
  enabled: true
  # v0.0.37+ required: it is the first release serving the /health and /ready
  # endpoints used by the startup/liveness/readiness probes in the
  # local-path-provisioner Deployment. Older images (e.g. v0.0.30) have no
  # health server, so the probes fail with "connection refused".
  image: rancher/local-path-provisioner:v0.0.37
  nodePath: /mnt/ssd-small/local
  # Tolerations for the per-node helper pods that create/remove the volume
  # directories. Must include the cluster's data/genai node taint so the
  # helper pod can run on the node hosting the PVC.
  helperPodTolerations:
    - key: node.kubernetes.io/not-ready
      operator: Exists
      effect: NoSchedule
    - key: node.kubernetes.io/disk-pressure
      operator: Exists
      effect: NoSchedule
    - key: opi5.cluster/role
      operator: Equal
      value: data
      effect: NoSchedule
  # Tolerations for the provisioner Deployment pod itself.
  tolerations:
    - key: opi5.cluster/role
      operator: Equal
      value: data
      effect: NoSchedule
  storageClass:
    name: local-path-ssd
    reclaimPolicy: Retain
    defaultClass: false

Manifests & templates

templates/http-route.yaml

Gateway API HTTPRoute exposing the service through the Istio gateway under opi5cluster.co.uk.

Show manifest
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: longhorn-httproute
  namespace: longhorn
  annotations:
    link.argocd.argoproj.io/external-link: "https://longhorn.opi5cluster.co.uk"
spec:
  parentRefs:
    - name: istio-gateway
      namespace: istio
      sectionName: websecure
  hostnames:
    - longhorn.opi5cluster.co.uk
  rules:
    - backendRefs:
        - name: longhorn-frontend
          port: 80

templates/storage-class-ssd-large.yaml

Longhorn StorageClass (ssd-large): replica count and replica placement tuning.

Show manifest
kind: StorageClass
apiVersion: storage.k8s.io/v1
metadata:
  name: longhorn-ssd-large
  annotations:
    longhorn.io/last-applied-configmap: |
      kind: StorageClass
      apiVersion: storage.k8s.io/v1
      metadata:
        name: longhorn-ssd-large
        annotations:
          storageclass.kubernetes.io/is-default-class: "false"
    storageclass.kubernetes.io/is-default-class: 'false'
provisioner: driver.longhorn.io
allowVolumeExpansion: true
reclaimPolicy: Retain
volumeBindingMode: Immediate
parameters:
  numberOfReplicas: "2"
  staleReplicaTimeout: "2880"
  fsType: "xfs"
  dataLocality: "best-effort"
  diskSelector: "ssd-large"

templates/storage-class-ssd-small.yaml

Longhorn StorageClass (ssd-small): replica count and replica placement tuning.

Show manifest
kind: StorageClass
apiVersion: storage.k8s.io/v1
metadata:
  name: longhorn-ssd-small
  annotations:
    longhorn.io/last-applied-configmap: |
      kind: StorageClass
      apiVersion: storage.k8s.io/v1
      metadata:
        name: longhorn-ssd-small
        annotations:
          storageclass.kubernetes.io/is-default-class: "true"
    storageclass.kubernetes.io/is-default-class: 'true'
provisioner: driver.longhorn.io
allowVolumeExpansion: true
reclaimPolicy: Retain
volumeBindingMode: Immediate
parameters:
  numberOfReplicas: "2"
  staleReplicaTimeout: "2880"
  fsType: "ext4"
  dataLocality: "best-effort"
  diskSelector: "ssd-small"

Trade-offs

Decision. Longhorn

Alternative. Rook/Ceph

Why. Rejected early (see the Lessons page): Ceph operational overhead and arm64 quirks are not worth it at this scale, while Longhorn is a UI away from understandable.

← Back to Platform & Infrastructure · All service groups