Decision. AutoMQ
Alternative. Redpanda or vanilla Kafka
Why. S3-backed log storage keeps the heavy state off small arm64 disks, and full Kafka API compatibility means no client changes.
Kafka with logs on S3-compatible storage: the event backbone every producer and consumer shares
AutoMQ is Kafka-API-compatible with tiered storage: log segments land on RustFS over the S3 API instead of filling broker disks. It is the event backbone the bots and services publish through.
Excerpt from argocd-apps/values.yaml in the argocd-apps chart,
with annotations added for this site:
automq:
application: true
project: cluster-services
# false: broker upgrades rebalance partitions; do them deliberately
autoSync: false
The full automq/values.yaml from the service's
own chart:
# RustFS bucket bootstrap Job (pre-install/pre-upgrade Helm hook).
# Values here are chart-local — the subchart cannot read parent values, so the
# endpoint/bucket names are repeated in automq.controller.extraConfig below.
bucketJob:
enabled: true
# rc CLI image built in the rustfs project (multi-arch, ENTRYPOINT ["rc"]).
image: registry.opi5cluster.co.uk/rustfs/cli:v0.1.31
# Canonical in-cluster RustFS S3 endpoint (same as argo-workflows).
endpoint: http://rustfs-service.rustfs.svc.cluster.local:9000
buckets:
- automq-data
- automq-ops
# External (LAN) exposure via traefik TCP/HTTP routing. DNS rewrites for
# both hostnames live in home-setup/adguard. Kafka rides a dedicated traefik
# entrypoint (port 9095) because plain TCP cannot share the HTTPS wildcard
# listener; the hostname only matters for DNS — clients connect by port.
external:
kafka:
enabled: true
hostname: automq.kafka.opi5cluster.co.uk
# LAN-facing port (redpanda gone; conventional Kafka port free again).
# The broker still BINDS the EXTERNAL listener on container port 9095 —
# 9092 inside the container belongs to the CLIENT listener — and the
# dedicated automq-external Service maps 9092 -> 9095.
port: 9092
schemaRegistry:
enabled: true
hostname: automq.schemaregistry.opi5cluster.co.uk
# Kafbat UI — web UI for the cluster (community successor of Provectus
# kafka-ui; multi-arch, incl. arm64). Unpinned: schedules on any untainted
# node. Exposed at https://automq.opi5cluster.co.uk (needs the AdGuard DNS rewrite
# in home-setup; no auth — internal LAN, same posture as the rustfs console).
ui:
enabled: true
image: ghcr.io/kafbat/kafka-ui:v1.5.0
resources:
requests:
cpu: 100m
memory: 512Mi
limits:
cpu: 500m
memory: 1Gi
# Confluent Schema Registry — reference implementation, multi-arch (arm64).
# Stores schemas in the compacted _schemas topic on AutoMQ. 7.9.x is the
# mature line paired with Kafka 3.x-era clients; single replica, KRaft mode
# (no ZooKeeper), internal-only Service (no HTTPRoute).
schemaRegistry:
enabled: true
image: confluentinc/cp-schema-registry:7.9.9
resources:
requests:
cpu: 100m
memory: 512Mi
limits:
cpu: "1"
memory: 1Gi
# ---------------------------------------------------------------------------
# Bitnami Kafka chart (aliased as "automq") running the AutoMQ image.
# All subchart values nest under the automq: alias.
# ---------------------------------------------------------------------------
automq:
global:
security:
# Required by the Bitnami chart to accept the non-Bitnami AutoMQ image.
allowInsecureImages: true
image:
registry: automqinc
repository: automq
tag: 1.7.3-bitnami
pullPolicy: IfNotPresent
# AutoMQ's S3Stream reads object-storage credentials from the standard AWS
# env vars. object-storage-secret is Reflector-replicated into this
# namespace from argo-workflows (the original), so the RustFS admin keys
# arrive automatically once the namespace exists — see README.
extraEnvVars:
- name: AWS_ACCESS_KEY_ID
valueFrom:
secretKeyRef:
name: object-storage-secret
key: access-key
- name: AWS_SECRET_ACCESS_KEY
valueFrom:
secretKeyRef:
name: object-storage-secret
key: secret-key
# Internal-only cluster behind traefik/AdGuard: no listener auth, matching
# the posture of rustfs/argo-workflows/redis. The Bitnami defaults are
# SASL_PLAINTEXT, so PLAINTEXT must be set explicitly on every listener.
# EXTERNAL (9095) is the LAN-facing listener advertised as
# automq.kafka.opi5cluster.co.uk:9095 — see external.kafka above.
listeners:
client:
protocol: PLAINTEXT
controller:
protocol: PLAINTEXT
interbroker:
protocol: PLAINTEXT
extraListeners:
- name: EXTERNAL
containerPort: 9095
protocol: PLAINTEXT
sslClientAuth: ""
# Fully-literal advertised listeners. The chart's init script rewrites its
# advertised-address-placeholder ONLY when this override is unset (guarded
# by `if not .Values.listeners.advertisedListeners` in scripts-configmap) —
# with an override, every host must be resolvable exactly as written.
# CLIENT advertises the shared Service DNS (stable, what in-cluster clients
# bootstrap against); INTERNAL advertises the pod-ordinal headless DNS
# (valid because controller.replicaCount is 1 — revisit if scaling, since
# each pod needs its own ordinal); EXTERNAL advertises the LAN hostname and
# the traefik entrypoint port clients dial (not the container bind port).
# CONTROLLER stays unadvertised (KRaft requirement).
advertisedListeners: "CLIENT://automq.automq.svc.cluster.local:9092,INTERNAL://automq-controller-0.automq-controller-headless.automq.svc.cluster.local:9094,EXTERNAL://automq.kafka.opi5cluster.co.uk:9092"
# The chart's NetworkPolicy only opens the client/interbroker/controller
# ports — the EXTERNAL listener port is included when externalAccess is
# enabled, which we don't use. Open the bind port (9095) explicitly so
# traefik can reach it.
networkPolicy:
extraIngress:
- ports:
- port: 9095
controller:
# Single combined controller+broker pod (AutoMQ quickstart topology).
# Scale up later by raising this number — never scale controllers down
# after the cluster has metadata.
replicaCount: 1
nodeSelector:
kubernetes.io/hostname: opi5-worker-5-genai
# opi5-worker-5-genai is tainted opi5.cluster/role=data:NoSchedule.
tolerations:
- key: opi5.cluster/role
operator: Equal
value: data
effect: NoSchedule
heapOpts: -Xmx1g -Xms1g -XX:MaxDirectMemorySize=1g
resources:
requests:
cpu: 500m
memory: 2Gi
# Small metadata disk only — all log data and the WAL live in RustFS.
persistence:
storageClass: longhorn-ssd-small
size: 5Gi
# Caches/bandwidth halved from AutoMQ's demo values for homelab scale.
# Bucket URIs must match bucketJob.endpoint/buckets above.
extraConfig: |
elasticstream.enable=true
s3.wal.cache.size=536870912
s3.block.cache.size=536870912
s3.stream.allocator.policy=POOLED_DIRECT
s3.network.baseline.bandwidth=52428800
s3.ops.buckets=1@s3://automq-ops?region=us-east-1&endpoint=http://rustfs-service.rustfs.svc.cluster.local:9000&pathStyle=true
s3.data.buckets=0@s3://automq-data?region=us-east-1&endpoint=http://rustfs-service.rustfs.svc.cluster.local:9000&pathStyle=true
s3.wal.path=0@s3://automq-data?region=us-east-1&endpoint=http://rustfs-service.rustfs.svc.cluster.local:9000&pathStyle=true
broker:
# Combined mode: controller pods serve the broker role. Raise this only to
# split dedicated broker pods out (broker IDs then start at 1000).
replicaCount: 0 templates/bucket-job.yaml Job that creates the S3 bucket prerequisites on RustFS before AutoMQ starts.
{{- if .Values.bucketJob.enabled }}
apiVersion: batch/v1
kind: Job
metadata:
name: {{ include "automq.bucketJobName" . }}
namespace: {{ .Release.Namespace }}
annotations:
# ArgoCD translates these to PreSync hooks; runs before the StatefulSet.
helm.sh/hook: pre-install,pre-upgrade
helm.sh/hook-weight: "-5"
helm.sh/hook-delete-policy: before-hook-creation
spec:
backoffLimit: 3
# Fails visibly (rather than hanging the sync) if the Reflector-copied
# secret or the RustFS service never becomes reachable.
activeDeadlineSeconds: 300
template:
metadata:
labels:
app: {{ include "automq.bucketJobName" . }}
spec:
restartPolicy: Never
containers:
- name: bucket-init
# rc image has ENTRYPOINT ["rc"]; the /bin/sh command replaces it.
image: {{ .Values.bucketJob.image }}
imagePullPolicy: IfNotPresent
env:
- name: ACCESS_KEY
valueFrom:
secretKeyRef:
name: object-storage-secret
key: access-key
- name: SECRET_KEY
valueFrom:
secretKeyRef:
name: object-storage-secret
key: secret-key
command:
- /bin/sh
- -c
- |
set -eu
rc alias set rustfs {{ .Values.bucketJob.endpoint }} "$ACCESS_KEY" "$SECRET_KEY"
for bucket in {{ .Values.bucketJob.buckets | join " " }}; do
if rc ls "rustfs/$bucket/" >/dev/null 2>&1; then
echo "bucket $bucket already exists"
else
rc mb "rustfs/$bucket"
fi
done
{{- end }} templates/kafka/kafka-external-route.yaml Dedicated LAN Service for the external listener plus the TCPRoute that publishes kafka :9092 through the Istio gateway.
{{- if .Values.external.kafka.enabled }}
apiVersion: v1
kind: Service
metadata:
name: automq-external-service
namespace: {{ .Release.Namespace }}
labels:
app.kubernetes.io/name: automq
app.kubernetes.io/instance: {{ .Release.Namespace }}
spec:
type: ClusterIP
# Same selector as the chart's shared client Service (combined mode).
selector:
app.kubernetes.io/name: automq
app.kubernetes.io/instance: {{ .Release.Namespace }}
app.kubernetes.io/component: controller-eligible
app.kubernetes.io/part-of: kafka
ports:
- name: tcp-external
port: {{ .Values.external.kafka.port }}
protocol: TCP
# Broker speaks Kafka on container port 9092; 9095 (legacy traefik-era
# bind port) answers nothing.
targetPort: 9092
---
apiVersion: gateway.networking.k8s.io/v1alpha2
kind: TCPRoute
metadata:
name: {{ .Release.Namespace }}-kafka-external
namespace: {{ .Release.Namespace }}
spec:
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: istio-gateway
namespace: istio
sectionName: kafka
rules:
- backendRefs:
- group: ''
kind: Service
name: automq-external-service
port: 9092
weight: 1
{{- end }} templates/kafka-ui.yaml Kafka UI rollout for inspecting AutoMQ topics and consumer lag.
{{- if .Values.ui.enabled }}
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: {{ .Release.Namespace }}-kafka-ui
namespace: {{ .Release.Namespace }}
labels:
app: {{ .Release.Namespace }}-kafka-ui
spec:
replicas: 1
strategy:
canary:
steps:
- setWeight: 25
- pause: {duration: 30s}
- setWeight: 50
- pause: {duration: 45s}
- setWeight: 100
selector:
matchLabels:
app: {{ .Release.Namespace }}-kafka-ui
template:
metadata:
labels:
app: {{ .Release.Namespace }}-kafka-ui
spec:
containers:
- name: kafka-ui
image: {{ .Values.ui.image }}
imagePullPolicy: IfNotPresent
env:
- name: KAFKA_CLUSTERS_0_NAME
value: automq
- name: KAFKA_CLUSTERS_0_BOOTSTRAPSERVERS
value: automq.{{ .Release.Namespace }}.svc.cluster.local:9092
{{- if .Values.schemaRegistry.enabled }}
- name: KAFKA_CLUSTERS_0_SCHEMAREGISTRY
value: http://{{ .Release.Namespace }}-schema-registry.{{ .Release.Namespace }}.svc.cluster.local:8081
{{- end }}
# TieredStopAtLevel=1: skip full JIT during startup — cuts cold
# start substantially on ARM; throughput penalty is irrelevant
# for an admin UI. egd avoids SecureRandom blocking in containers.
- name: JAVA_OPTS
value: -Xmx512m -XX:TieredStopAtLevel=1 -Djava.security.egd=file:/dev/./urandom
ports:
- name: http
containerPort: 8080
resources: {{- toYaml .Values.ui.resources | nindent 12 }}
# Kafbat health endpoint is /actuator/health (per their own compose
# healthchecks); /healthcheck 404s and crash-looped the pod via
# liveness. startupProbe absorbs the Java cold start; timeoutSeconds
# is raised because the default 1s produces false failures while
# Spring is still initializing after Jetty binds.
startupProbe:
httpGet:
path: /actuator/health
port: http
periodSeconds: 5
timeoutSeconds: 5
failureThreshold: 60
livenessProbe:
httpGet:
path: /actuator/health
port: http
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 5
readinessProbe:
httpGet:
path: /actuator/health
port: http
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 6
---
apiVersion: v1
kind: Service
metadata:
name: {{ .Release.Namespace }}-kafka-ui
namespace: {{ .Release.Namespace }}
labels:
app: {{ .Release.Namespace }}-kafka-ui
spec:
type: ClusterIP
selector:
app: {{ .Release.Namespace }}-kafka-ui
ports:
- name: http
port: 80
targetPort: http
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: {{ .Release.Namespace }}-kafka-ui
namespace: {{ .Release.Namespace }}
annotations:
link.argocd.argoproj.io/external-link: "https://automq.opi5cluster.co.uk"
spec:
parentRefs:
- name: istio-gateway
namespace: istio
sectionName: websecure
hostnames:
- automq.opi5cluster.co.uk
rules:
- backendRefs:
- name: {{ .Release.Namespace }}-kafka-ui
port: 80
{{- end }} templates/schema-registry.yaml Schema Registry rollout for Avro payloads on AutoMQ.
{{- if .Values.schemaRegistry.enabled }}
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: {{ .Release.Namespace }}-schema-registry
namespace: {{ .Release.Namespace }}
labels:
app: {{ .Release.Namespace }}-schema-registry
spec:
replicas: 1
strategy:
canary:
steps:
- setWeight: 25
- pause: {duration: 30s}
- setWeight: 50
- pause: {duration: 45s}
- setWeight: 100
selector:
matchLabels:
app: {{ .Release.Namespace }}-schema-registry
template:
metadata:
labels:
app: {{ .Release.Namespace }}-schema-registry
spec:
containers:
- name: schema-registry
image: {{ .Values.schemaRegistry.image }}
imagePullPolicy: IfNotPresent
env:
# KRaft mode: schema state lives in the compacted _schemas topic
# on AutoMQ — no ZooKeeper, no PVC. NOTE: bare host:port only —
# a "plaintext://" prefix is parsed as a listener name by
# Confluent SR and rejected ("No supported Kafka endpoints").
- name: SCHEMA_REGISTRY_KAFKASTORE_BOOTSTRAP_SERVERS
value: automq.{{ .Release.Namespace }}.svc.cluster.local:9092
- name: SCHEMA_REGISTRY_KAFKASTORE_SECURITY_PROTOCOL
value: PLAINTEXT
- name: SCHEMA_REGISTRY_KAFKASTORE_TOPIC
value: _schemas
# Must match broker count — the default (3) kills topic creation
# on this single-broker cluster. Raise alongside controller.replicaCount.
- name: SCHEMA_REGISTRY_KAFKASTORE_TOPIC_REPLICATION_FACTOR
value: "1"
- name: SCHEMA_REGISTRY_HOST_NAME
valueFrom:
fieldRef:
fieldPath: metadata.name
- name: SCHEMA_REGISTRY_LISTENERS
value: http://0.0.0.0:8081
- name: SCHEMA_REGISTRY_SCHEMA_REGISTRY_GROUP_ID
value: schema-registry
- name: SCHEMA_REGISTRY_HEAP_OPTS
value: -Xms256m -Xmx512m
ports:
- name: http
containerPort: 8081
resources: {{- toYaml .Values.schemaRegistry.resources | nindent 12 }}
livenessProbe:
httpGet:
path: /subjects
port: http
periodSeconds: 30
failureThreshold: 6
readinessProbe:
httpGet:
path: /subjects
port: http
periodSeconds: 10
failureThreshold: 18
---
apiVersion: v1
kind: Service
metadata:
name: {{ .Release.Namespace }}-schema-registry
namespace: {{ .Release.Namespace }}
labels:
app: {{ .Release.Namespace }}-schema-registry
spec:
type: ClusterIP
selector:
app: {{ .Release.Namespace }}-schema-registry
ports:
- name: http
port: 8081
targetPort: http
{{- end }}
{{- if and .Values.schemaRegistry.enabled .Values.external.schemaRegistry.enabled }}
---
# Schema Registry is plain HTTP(S) — standard house HTTPRoute on the shared
# websecure listener, wildcard *.opi5cluster.co.uk TLS. External clients use
# https://automq.schemaregistry.opi5cluster.co.uk; in-cluster consumers keep the
# Service DNS (automq-schema-registry:8081).
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: {{ .Release.Namespace }}-schema-registry
namespace: {{ .Release.Namespace }}
annotations:
link.argocd.argoproj.io/external-link: "http://{{ .Values.external.schemaRegistry.hostname }}"
spec:
parentRefs:
- name: istio-gateway
namespace: istio
sectionName: web
hostnames:
- {{ .Values.external.schemaRegistry.hostname }}
rules:
- backendRefs:
- name: {{ .Release.Namespace }}-schema-registry
port: 8081
{{- end }} Decision. AutoMQ
Alternative. Redpanda or vanilla Kafka
Why. S3-backed log storage keeps the heavy state off small arm64 disks, and full Kafka API compatibility means no client changes.