Experiments & Lessons

What worked, what broke, and what's next.

The opi5-cluster is an experiment that never really ends. What started as a few ARM boards has grown into a mixed ARM/x86 Kubernetes platform running databases, data infrastructure, applications, and AI workloads. The goal isn't to build a perfect cluster. It's to learn by building, breaking, rebuilding, and progressively applying production-grade practices. Some experiments become permanent capabilities, others become useful lessons.

What worked?

  • Keeping the name: Calling a mixed ARM/x86 fleet opi5-cluster might sound wrong, but it is a reminder that clusters grow organically. The name doesn't necessarily need to be accurate to be good.
  • GitOps almost from day one: ArgoCD and an app-of-apps pattern give the cluster a consistent, declarative deployment mode, one pattern, and every application accounted for.
  • Semi Automated node provisioning: DietPi plus committed first-boot automation means any node can be rebuilt in less than a hour.
  • Running DNS outside the cluster: The two NanoPi NEO3 appliances keep household DNS working even when the Kubernetes cluster is having a bad day. Not every dependency needs to live inside Kubernetes.
  • Choosing simplicity over unnecessary complexity: Experiments with Rook/Ceph, Druid, Kafka, Redpanda and several workflow/technology platforms helped establish that the most sophisticated solution isn't always the best fit. Longhorn, AutoMQ and Argo Workflows ultimately provided better fits for this environment.
  • Service mesh by migration, not big-bang: Replacing Traefik with an Istio ambient mesh moved one route at a time: dual parentRefs, a curl --resolve gate, per-host DNS flips, a burn-in period, then stripping the old parent. All 32 routes moved with zero user-visible outages, every step was one DNS flip away from rollback, and mTLS with SPIFFE identities between workloads came with the ambient enrolment, not as extra work. Traefik was only decommissioned once its listeners provably carried no traffic.

What broke?

  • NVMe/SSD matters more than you think: One of the lessons I learned the hard way is that where K3s stores its data matters.After a few outages involving boot SD cards, it became clear that the K3s data directory had no business living on the same fragile storage the system boots from. Moving it to local NVMe/SSD storage made a significant difference in both reliability and recovery.
  • Not everything belongs on Longhorn: Distributed storage isn't automatically better storage. PostgreSQL is deployed as one primary and two replicas, so there is already redundancy at the database layer. Putting those databases on Longhorn PVs would effectively add another layer of replication and distribution on top. That means Longhorn ends up moving and rebalancing data that PostgreSQL is already replicating itself, a lot of unnecessary work for very little benefit. For PostgreSQL, I now prefer local storage directly on the nodes, letting PostgreSQL handle replication and keeping Longhorn out of the equation.
  • More redundancy is not always more resilience: Several experiments reinforced that every additional distributed layer brings its own complexity and failure modes. Resilience comes from putting redundancy at the right layer, not simply adding more of it.
  • Kubernetes-native doesn't always mean Kubernetes-required: Some workloads benefit enormously from Kubernetes; others are better served by simpler infrastructure. The cluster has become more reliable as I have become more selective about what actually needs to be distributed or orchestrated.
  • The newest layer is usually not the culprit: During the Istio rollout, cross-node mesh traffic kept dying in a way that looked exactly like half-enrolled workloads. Two days and five gated experiments later, the cause turned out to be a kube-router NetworkPolicy silently rejecting the mesh's HBONE port, with REJECT rules even impersonating a kernel connection-refused. The fix was one firewall rule. Read the live evidence, sockets and packet captures, before blaming the component you just installed.

What I've learned?

The biggest lesson has been that production-grade engineering is mostly about trade-offs. Hardware matters, storage strategy matters, failure domains matter, observability matters...and sometimes the right answer is to not add another layer. My opi5-cluster provides a safe environment to learn these lessons through real failures rather than theory, and every experiment has made the next iteration a little more deliberate.

What could be next?

The cluster is continuously evolving, not just by adding more services, but by progressively improving its architecture, resilience, security, and operational maturity. Each next step is an opportunity to move closer to production-grade practices while exploring a new area of cloud-native engineering.

  • Build a highly available control plane: Evolve from a single control-plane node to three K3s servers backed by an HA etcd cluster, gaining more hands-on experience with quorum, failure recovery, control-plane resilience, and production-grade Kubernetes architecture.
  • Harden persistent storage for stateful workloads: Move the media stack away from node-pinned hostPath mounts toward replicated storage, exploring the trade-offs between performance, availability, data locality, and operational simplicity for stateful workloads.
  • Expand the local AI platform: Continue experimenting with locally hosted models on the dedicated genai node, exploring model serving, resource allocation, inference performance, GPU acceleration, and the operational challenges of running AI workloads alongside a Kubernetes platform.
  • Introduce Policy-as-Code: Evaluate admission controllers such as Kyverno to enforce security, configuration, and operational policies across workloads, gaining practical experience with governance and preventative controls in Kubernetes.
  • Expand Progressive Delivery: Build on the existing use of Argo Rollouts beyond simple weight-based canaries, exploring automated analysis, health-based promotion, rollback strategies, and more sophisticated deployment patterns. The aim is to make progressive delivery a consistent production practice rather than something applied only to selected services.
  • Introduce Centralised Identity and Access Management: Evaluate open-source IAM platforms such as Keycloak to provide a central authentication and authorisation layer based on standards such as OpenID Connect, OAuth 2.0, and SAML. This would also provide a practical foundation for exploring SSO and identity-aware access across cluster services.