Decision. 30-day retention on local storage
Alternative. A remote long-term metrics store (Thanos or Mimir)
Why. Lab-scale dashboards do not justify a second storage system; manifests are covered by Velero anyway.
Scrapes metrics from every workload and graphs them in Grafana, 30-day retention
The monitoring stack is a single Application: Prometheus scrapes the cluster, and Grafana serves dashboards for nodes, pods, storage, and the edge. Retention is 30 days, sized against Longhorn capacity.
Excerpt from argocd-apps/values.yaml in the argocd-apps chart,
with annotations added for this site:
monitoring:
# true
application: true
project: cluster-services
# false
autoSync: false
The full monitoring/values.yaml from the service's
own chart:
# ─────────────────────────────────────────────────────────────────────
# Centralized monitoring: Grafana Alloy (DaemonSet) → Prometheus + Loki,
# dashboards in Grafana. Self-hosted; no external SaaS, no CRDs.
# ─────────────────────────────────────────────────────────────────────
# ── Cluster-wide metadata attached to every metric / log ─────────────
cluster:
name: opi5-cluster
environment: prod
# ── Static scrape targets for the cluster services ───────────────────
# Addresses are in-cluster Service DNS names. Each entry was verified
# against a rendered upstream chart (not guessed). If a service name
# ever changes, fix it here.
#
# Live verification: kubectl get svc -A
#
# Some targets are only scraped once the corresponding chart change has
# been applied (see the per-chart "monitoring" commits):
# * argo-cd / argo-workflows / argo-rollouts — metrics were off by
# default and have been enabled in their values.
# * vault — telemetry stanza added (path /v1/sys/metrics).
# * openvino — metrics enabled on the REST port (`--metrics_enable`),
# scraped from the Service's :8000/metrics.
metricsScrapeTargets:
- name: istio-gateway
job: istio-gateway
url: istio-gateway-istio.istio.svc.cluster.local:15090
namespace: istio
component: gateway
- name: argocd_application_controller
job: argocd/application-controller
url: argocd-application-controller-metrics.argocd.svc.cluster.local:8082
namespace: argocd
component: controller
- name: argocd_server
job: argocd/server
url: argocd-server-metrics.argocd.svc.cluster.local:8083
namespace: argocd
component: server
- name: argocd_repo_server
job: argocd/repo-server
url: argocd-repo-server-metrics.argocd.svc.cluster.local:8084
namespace: argocd
component: repo-server
- name: argocd_applicationset_controller
job: argocd/applicationset-controller
url: argocd-applicationset-controller-metrics.argocd.svc.cluster.local:8080
namespace: argocd
component: applicationset-controller
- name: argo_rollouts_controller
job: argo-rollouts/controller
url: argo-rollouts-metrics.argo-rollouts.svc.cluster.local:8090
namespace: argo-rollouts
component: controller
- name: argo_workflows_controller
job: argo-workflows/controller
url: argo-workflows-controller.argo-workflows.svc.cluster.local:8080
namespace: argo-workflows
component: controller
- name: argo_workflows_server
job: argo-workflows/server
url: argo-workflows-server.argo-workflows.svc.cluster.local:2746
namespace: argo-workflows
component: server
- name: vault
job: vault
url: vault.vault.svc.cluster.local:8200
namespace: vault
component: server
path: /v1/sys/metrics?format=prometheus
- name: longhorn_manager
job: longhorn/manager
url: longhorn-backend.longhorn-system.svc.cluster.local:9500
namespace: longhorn-system
component: manager
- name: openvino_model_server
job: openvino/model-server
url: openvino-service.openvino.svc.cluster.local:8000
namespace: openvino
component: model-server
path: /metrics
- name: zot
job: zot/registry
url: zot.zot.svc.cluster.local:5000
namespace: zot
component: registry
path: /metrics
# ── Grafana Alloy (upstream chart) ───────────────────────────────────
# Note: the upstream alloy chart reads its runtime settings from
# `.Values.alloy` (and root-level `crds`/`controller`). Because Helm maps
# the parent's `alloy:` block to the subchart root, runtime settings need
# the nested `alloy.alloy:` key.
alloy:
alloy:
# We render our own ConfigMap (templates/alloy-config.yaml) so the
# River config is generated from the values above. The ConfigMap name
# defaults to the chart-computed fullname (`{release}-alloy`), so it
# matches whatever the DaemonSet expects regardless of release name.
configMap:
create: false
mounts:
# Logs are streamed via the Kubernetes API (loki.source.kubernetes),
# not tailed from the node filesystem — no host mounts required.
varlog: false
dockercontainers: false
storagePath: /tmp/alloy
enableReporting: false
# Cluster mode: with one Alloy pod per node, every pod was scraping the
# full cluster-wide target set, so Prometheus rejected all but the first
# copy of each series ("out of order sample from remote write").
# Clustering hash-shards the cluster-wide scrape targets so exactly one
# member owns each target; node-local scrapes (kubelet/cAdvisor and
# node-exporter pods) are instead scoped to the local node in the
# Alloy config (see templates/alloy-config.yaml), where each node has
# exactly one writer by construction.
clustering:
enabled: true
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 512Mi
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop:
- ALL
seccompProfile:
type: RuntimeDefault
# The chart's default `crds` subchart installs the
# monitoring.grafana.com/podlogs CRD — we don't use podlogs, so keep
# the cluster CRD-free.
crds:
create: false
# Alloy controller: DaemonSet, one pod per node.
controller:
type: daemonset
# Run on every worker node, including the tainted data/LLM nodes.
tolerations:
- key: opi5.cluster/role
operator: Equal
value: data
effect: NoSchedule
nodeSelector: {}
# ── kube-state-metrics (upstream chart) ──────────────────────────────
kube-state-metrics:
resources:
requests:
cpu: 20m
memory: 80Mi
limits:
cpu: 100m
memory: 200Mi
# ── Prometheus (upstream chart) ──────────────────────────────────────
# Pure metrics store + rule engine. Alloy is the only scraper and pushes
# via remote_write; Prometheus itself has no scrape configs to maintain.
# Only the server workload is deployed — Alertmanager, node-exporter,
# pushgateway and the bundled kube-state-metrics are all disabled.
prometheus:
alertmanager:
enabled: false
kube-state-metrics:
enabled: false
# node-exporter provides node filesystem/network/load metrics for the
# Host Overview dashboard. Alloy discovers its pods via the
# prometheus.io/scrape annotations below (annotation-based discovery),
# so no extra scrape config is needed.
prometheus-node-exporter:
enabled: true
podAnnotations:
prometheus.io/scrape: "true"
prometheus.io/port: "19100"
# hostNetwork pods bind the node's port directly. Traefik's metrics
# entrypoint holds host 9100 (traefik chart default) on opi5-worker-2,
# which kept this node's node-exporter Pending — so listen on 19100
# instead. service.port drives both the bind address and containerPort.
service:
port: 19100
targetPort: 19100
# DaemonSet must cover every worker node for Host Overview. The broad
# operator:Exists toleration is the chart default (any NoSchedule
# taint); the explicit data-role toleration below documents the
# intent to also run on the tainted data nodes.
tolerations:
- effect: NoSchedule
operator: Exists
- key: opi5.cluster/role
operator: Equal
value: data
effect: NoSchedule
prometheus-pushgateway:
enabled: false
server:
retention: 30d
# Accept remote_write from Alloy. Prometheus v3.13+ reverted to the
# --web.enable-remote-write-receiver flag (the remote-write-receiver
# entry was removed from --enable-feature), and the bundled
# config-reloader needs --web.enable-lifecycle to trigger reloads
# (403 otherwise).
extraFlags:
- web.enable-remote-write-receiver
- web.enable-lifecycle
service:
servicePort: 9090
persistentVolume:
size: 20Gi
# ── Loki (upstream chart) ────────────────────────────────────────────
# Single-binary mode (SSD/SimpleScalable is deprecated and removed in
# Loki 4). Filesystem storage on a Longhorn PVC — no object store needed.
# Note: `loki:` here maps to the subchart root; the Loki config sections
# are under `loki.loki`.
loki:
deploymentMode: SingleBinary
singleBinary:
replicas: 1
persistence:
size: 20Gi
# Zero out the SimpleScalable targets (defaults are non-zero).
write:
replicas: 0
read:
replicas: 0
backend:
replicas: 0
# Disable optional components we don't need.
gateway:
enabled: false
lokiCanary:
enabled: false
test:
enabled: false
monitoring:
selfMonitoring:
enabled: false
lokiCanary:
enabled: false
chunksCache:
enabled: false
resultsCache:
enabled: false
loki:
# Single-tenant cluster: the chart default (auth_enabled: true) makes
# Loki 401 every request without an X-Scope-OrgID header, which breaks
# both Alloy's log pushes and Grafana's queries.
auth_enabled: false
commonConfig:
# Single-binary single-replica: the chart default replication_factor
# of 3 leaves the ring with "too many unhealthy instances" and every
# query fails with 500.
replication_factor: 1
storage:
type: filesystem
schemaConfig:
configs:
- from: "2024-01-01"
store: tsdb
object_store: filesystem
schema: v13
index:
prefix: index_
period: 24h
# 14-day log retention. Loki 3.x requires delete_request_store when
# retention is enabled — "filesystem" names the local store.
limits_config:
retention_period: 336h
# Enable the log-volume endpoint so Grafana Explore can render the
# volume-over-time histogram.
volume_enabled: true
compactor:
retention_enabled: true
delete_request_store: filesystem
# ── Grafana (upstream chart) ─────────────────────────────────────────
# Admin credentials and the security secret_key come from Vault via ESO
# (Vault kv path `grafana` with properties admin_email, admin_password,
# admin_user, secret_key) — synced into the `monitoring-grafana-admin`
# secret by templates/grafana-external-secret.yaml. Keep the secret name
# in sync with that template's target.name.
grafana:
admin:
existingSecret: monitoring-grafana-admin
userKey: admin_user
passwordKey: admin_password
# [security] secret_key — Grafana reads GF_SECURITY_SECRET_KEY from env.
envValueFrom:
GF_SECURITY_SECRET_KEY:
secretKeyRef:
name: monitoring-grafana-admin
key: secret_key
# The chart default readinessProbe has no initialDelaySeconds, so it
# flaps "connection refused" while Grafana binds :3000 on every start
# (slower on a fresh Longhorn PVC + first-boot migrations).
readinessProbe:
initialDelaySeconds: 20
sidecar:
datasources:
enabled: true
label: grafana_datasource
# Load dashboards-as-code from ConfigMaps labelled
# `grafana_dashboard: "1"` (see templates/grafana-dashboards/).
dashboards:
enabled: true
label: grafana_dashboard
# Absolute path (the chart default is /tmp/dashboards). A relative
# value makes the sidecar resolve it against /app (its WORKDIR) and
# fail to write, crash-looping grafana-sc-dashboard.
folder: /Hosts
# Dashboards/datasources are provisioned via ConfigMap, but the SQLite
# DB (annotations, alert rules, orgs, folders) must survive pod
# restarts, so persist it on a Longhorn volume.
persistence:
enabled: true
size: 5Gi
storageClassName: longhorn
service:
port: 80
ingress:
enabled: false templates/grafana-datasources.yaml Grafana datasource provisioning (Prometheus, Loki).
apiVersion: v1
kind: ConfigMap
metadata:
name: {{ include "app.fullname" . }}-grafana-datasources
namespace: {{ .Release.Namespace }}
labels:
{{- include "app.labels" . | nindent 4 }}
# Picked up by the Grafana chart's datasource sidecar.
grafana_datasource: "1"
data:
datasources.yaml: |
apiVersion: 1
datasources:
- name: Prometheus
uid: prometheus
type: prometheus
access: proxy
url: http://{{ include "app.fullname" . }}-prometheus-server.{{ .Release.Namespace }}.svc.cluster.local:9090
isDefault: true
- name: Loki
uid: loki
type: loki
access: proxy
url: http://{{ include "app.fullname" . }}-loki.{{ .Release.Namespace }}.svc.cluster.local:3100 templates/grafana-external-secret.yaml ExternalSecret for Grafana admin and OAuth credentials.
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
name: {{ include "app.fullname" . }}-grafana-admin
namespace: {{ .Release.Namespace }}
labels:
{{- include "app.labels" . | nindent 4 }}
app.kubernetes.io/component: secrets
spec:
refreshInterval: 1h
secretStoreRef:
kind: ClusterSecretStore
name: vault-cluster-secret-store
target:
name: monitoring-grafana-admin
creationPolicy: Owner
data:
- secretKey: admin_email
remoteRef:
key: grafana
property: admin_email
- secretKey: admin_password
remoteRef:
key: grafana
property: admin_password
- secretKey: admin_user
remoteRef:
key: grafana
property: admin_user
- secretKey: secret_key
remoteRef:
key: grafana
property: secret_key templates/grafana-httproute.yaml HTTPRoute for the Grafana UI.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: {{ include "app.fullname" . }}-grafana-httproute
namespace: {{ .Release.Namespace }}
annotations:
link.argocd.argoproj.io/external-link: "https://grafana.opi5cluster.co.uk"
spec:
parentRefs:
- name: istio-gateway
namespace: istio
sectionName: websecure
hostnames:
- grafana.opi5cluster.co.uk
rules:
- backendRefs:
- name: {{ include "app.fullname" . }}-grafana
port: 80 Decision. 30-day retention on local storage
Alternative. A remote long-term metrics store (Thanos or Mimir)
Why. Lab-scale dashboards do not justify a second storage system; manifests are covered by Velero anyway.