Prometheus + Grafana

Scrapes metrics from every workload and graphs them in Grafana, 30-day retention

The monitoring stack is a single Application: Prometheus scrapes the cluster, and Grafana serves dashboards for nodes, pods, storage, and the edge. Retention is 30 days, sized against Longhorn capacity.

ArgoCD configuration

Excerpt from argocd-apps/values.yaml in the argocd-apps chart, with annotations added for this site:

monitoring:
  # true
  application: true
  project: cluster-services
  # false
  autoSync: false

Chart values

The full monitoring/values.yaml from the service's own chart:

# ─────────────────────────────────────────────────────────────────────
# Centralized monitoring: Grafana Alloy (DaemonSet) → Prometheus + Loki,
# dashboards in Grafana. Self-hosted; no external SaaS, no CRDs.
# ─────────────────────────────────────────────────────────────────────

# ── Cluster-wide metadata attached to every metric / log ─────────────
cluster:
  name: opi5-cluster
  environment: prod

# ── Static scrape targets for the cluster services ───────────────────
# Addresses are in-cluster Service DNS names. Each entry was verified
# against a rendered upstream chart (not guessed). If a service name
# ever changes, fix it here.
#
# Live verification:   kubectl get svc -A
#
# Some targets are only scraped once the corresponding chart change has
# been applied (see the per-chart "monitoring" commits):
#   * argo-cd / argo-workflows / argo-rollouts — metrics were off by
#     default and have been enabled in their values.
#   * vault — telemetry stanza added (path /v1/sys/metrics).
#   * openvino — metrics enabled on the REST port (`--metrics_enable`),
#     scraped from the Service's :8000/metrics.
metricsScrapeTargets:
  - name: istio-gateway
    job: istio-gateway
    url: istio-gateway-istio.istio.svc.cluster.local:15090
    namespace: istio
    component: gateway
  - name: argocd_application_controller
    job: argocd/application-controller
    url: argocd-application-controller-metrics.argocd.svc.cluster.local:8082
    namespace: argocd
    component: controller
  - name: argocd_server
    job: argocd/server
    url: argocd-server-metrics.argocd.svc.cluster.local:8083
    namespace: argocd
    component: server
  - name: argocd_repo_server
    job: argocd/repo-server
    url: argocd-repo-server-metrics.argocd.svc.cluster.local:8084
    namespace: argocd
    component: repo-server
  - name: argocd_applicationset_controller
    job: argocd/applicationset-controller
    url: argocd-applicationset-controller-metrics.argocd.svc.cluster.local:8080
    namespace: argocd
    component: applicationset-controller
  - name: argo_rollouts_controller
    job: argo-rollouts/controller
    url: argo-rollouts-metrics.argo-rollouts.svc.cluster.local:8090
    namespace: argo-rollouts
    component: controller
  - name: argo_workflows_controller
    job: argo-workflows/controller
    url: argo-workflows-controller.argo-workflows.svc.cluster.local:8080
    namespace: argo-workflows
    component: controller
  - name: argo_workflows_server
    job: argo-workflows/server
    url: argo-workflows-server.argo-workflows.svc.cluster.local:2746
    namespace: argo-workflows
    component: server
  - name: vault
    job: vault
    url: vault.vault.svc.cluster.local:8200
    namespace: vault
    component: server
    path: /v1/sys/metrics?format=prometheus
  - name: longhorn_manager
    job: longhorn/manager
    url: longhorn-backend.longhorn-system.svc.cluster.local:9500
    namespace: longhorn-system
    component: manager
  - name: openvino_model_server
    job: openvino/model-server
    url: openvino-service.openvino.svc.cluster.local:8000
    namespace: openvino
    component: model-server
    path: /metrics
  - name: zot
    job: zot/registry
    url: zot.zot.svc.cluster.local:5000
    namespace: zot
    component: registry
    path: /metrics

# ── Grafana Alloy (upstream chart) ───────────────────────────────────
# Note: the upstream alloy chart reads its runtime settings from
# `.Values.alloy` (and root-level `crds`/`controller`). Because Helm maps
# the parent's `alloy:` block to the subchart root, runtime settings need
# the nested `alloy.alloy:` key.
alloy:
  alloy:
    # We render our own ConfigMap (templates/alloy-config.yaml) so the
    # River config is generated from the values above. The ConfigMap name
    # defaults to the chart-computed fullname (`{release}-alloy`), so it
    # matches whatever the DaemonSet expects regardless of release name.
    configMap:
      create: false
    mounts:
      # Logs are streamed via the Kubernetes API (loki.source.kubernetes),
      # not tailed from the node filesystem — no host mounts required.
      varlog: false
      dockercontainers: false
    storagePath: /tmp/alloy
    enableReporting: false
    # Cluster mode: with one Alloy pod per node, every pod was scraping the
    # full cluster-wide target set, so Prometheus rejected all but the first
    # copy of each series ("out of order sample from remote write").
    # Clustering hash-shards the cluster-wide scrape targets so exactly one
    # member owns each target; node-local scrapes (kubelet/cAdvisor and
    # node-exporter pods) are instead scoped to the local node in the
    # Alloy config (see templates/alloy-config.yaml), where each node has
    # exactly one writer by construction.
    clustering:
      enabled: true
    resources:
      requests:
        cpu: 100m
        memory: 128Mi
      limits:
        cpu: 500m
        memory: 512Mi
    securityContext:
      allowPrivilegeEscalation: false
      capabilities:
        drop:
          - ALL
      seccompProfile:
        type: RuntimeDefault

  # The chart's default `crds` subchart installs the
  # monitoring.grafana.com/podlogs CRD — we don't use podlogs, so keep
  # the cluster CRD-free.
  crds:
    create: false

  # Alloy controller: DaemonSet, one pod per node.
  controller:
    type: daemonset
    # Run on every worker node, including the tainted data/LLM nodes.
    tolerations:
      - key: opi5.cluster/role
        operator: Equal
        value: data
        effect: NoSchedule
    nodeSelector: {}

# ── kube-state-metrics (upstream chart) ──────────────────────────────
kube-state-metrics:
  resources:
    requests:
      cpu: 20m
      memory: 80Mi
    limits:
      cpu: 100m
      memory: 200Mi

# ── Prometheus (upstream chart) ──────────────────────────────────────
# Pure metrics store + rule engine. Alloy is the only scraper and pushes
# via remote_write; Prometheus itself has no scrape configs to maintain.
# Only the server workload is deployed — Alertmanager, node-exporter,
# pushgateway and the bundled kube-state-metrics are all disabled.
prometheus:
  alertmanager:
    enabled: false
  kube-state-metrics:
    enabled: false
  # node-exporter provides node filesystem/network/load metrics for the
  # Host Overview dashboard. Alloy discovers its pods via the
  # prometheus.io/scrape annotations below (annotation-based discovery),
  # so no extra scrape config is needed.
  prometheus-node-exporter:
    enabled: true
    podAnnotations:
      prometheus.io/scrape: "true"
      prometheus.io/port: "19100"
    # hostNetwork pods bind the node's port directly. Traefik's metrics
    # entrypoint holds host 9100 (traefik chart default) on opi5-worker-2,
    # which kept this node's node-exporter Pending — so listen on 19100
    # instead. service.port drives both the bind address and containerPort.
    service:
      port: 19100
      targetPort: 19100
    # DaemonSet must cover every worker node for Host Overview. The broad
    # operator:Exists toleration is the chart default (any NoSchedule
    # taint); the explicit data-role toleration below documents the
    # intent to also run on the tainted data nodes.
    tolerations:
      - effect: NoSchedule
        operator: Exists
      - key: opi5.cluster/role
        operator: Equal
        value: data
        effect: NoSchedule
  prometheus-pushgateway:
    enabled: false
  server:
    retention: 30d
    # Accept remote_write from Alloy. Prometheus v3.13+ reverted to the
    # --web.enable-remote-write-receiver flag (the remote-write-receiver
    # entry was removed from --enable-feature), and the bundled
    # config-reloader needs --web.enable-lifecycle to trigger reloads
    # (403 otherwise).
    extraFlags:
      - web.enable-remote-write-receiver
      - web.enable-lifecycle
    service:
      servicePort: 9090
    persistentVolume:
      size: 20Gi

# ── Loki (upstream chart) ────────────────────────────────────────────
# Single-binary mode (SSD/SimpleScalable is deprecated and removed in
# Loki 4). Filesystem storage on a Longhorn PVC — no object store needed.
# Note: `loki:` here maps to the subchart root; the Loki config sections
# are under `loki.loki`.
loki:
  deploymentMode: SingleBinary
  singleBinary:
    replicas: 1
    persistence:
      size: 20Gi
  # Zero out the SimpleScalable targets (defaults are non-zero).
  write:
    replicas: 0
  read:
    replicas: 0
  backend:
    replicas: 0
  # Disable optional components we don't need.
  gateway:
    enabled: false
  lokiCanary:
    enabled: false
  test:
    enabled: false
  monitoring:
    selfMonitoring:
      enabled: false
    lokiCanary:
      enabled: false
  chunksCache:
    enabled: false
  resultsCache:
    enabled: false
  loki:
    # Single-tenant cluster: the chart default (auth_enabled: true) makes
    # Loki 401 every request without an X-Scope-OrgID header, which breaks
    # both Alloy's log pushes and Grafana's queries.
    auth_enabled: false
    commonConfig:
      # Single-binary single-replica: the chart default replication_factor
      # of 3 leaves the ring with "too many unhealthy instances" and every
      # query fails with 500.
      replication_factor: 1
    storage:
      type: filesystem
    schemaConfig:
      configs:
        - from: "2024-01-01"
          store: tsdb
          object_store: filesystem
          schema: v13
          index:
            prefix: index_
            period: 24h
    # 14-day log retention. Loki 3.x requires delete_request_store when
    # retention is enabled — "filesystem" names the local store.
    limits_config:
      retention_period: 336h
      # Enable the log-volume endpoint so Grafana Explore can render the
      # volume-over-time histogram.
      volume_enabled: true
    compactor:
      retention_enabled: true
      delete_request_store: filesystem

# ── Grafana (upstream chart) ─────────────────────────────────────────
# Admin credentials and the security secret_key come from Vault via ESO
# (Vault kv path `grafana` with properties admin_email, admin_password,
# admin_user, secret_key) — synced into the `monitoring-grafana-admin`
# secret by templates/grafana-external-secret.yaml. Keep the secret name
# in sync with that template's target.name.
grafana:
  admin:
    existingSecret: monitoring-grafana-admin
    userKey: admin_user
    passwordKey: admin_password
  # [security] secret_key — Grafana reads GF_SECURITY_SECRET_KEY from env.
  envValueFrom:
    GF_SECURITY_SECRET_KEY:
      secretKeyRef:
        name: monitoring-grafana-admin
        key: secret_key
  # The chart default readinessProbe has no initialDelaySeconds, so it
  # flaps "connection refused" while Grafana binds :3000 on every start
  # (slower on a fresh Longhorn PVC + first-boot migrations).
  readinessProbe:
    initialDelaySeconds: 20
  sidecar:
    datasources:
      enabled: true
      label: grafana_datasource
    # Load dashboards-as-code from ConfigMaps labelled
    # `grafana_dashboard: "1"` (see templates/grafana-dashboards/).
    dashboards:
      enabled: true
      label: grafana_dashboard
      # Absolute path (the chart default is /tmp/dashboards). A relative
      # value makes the sidecar resolve it against /app (its WORKDIR) and
      # fail to write, crash-looping grafana-sc-dashboard.
      folder: /Hosts
  # Dashboards/datasources are provisioned via ConfigMap, but the SQLite
  # DB (annotations, alert rules, orgs, folders) must survive pod
  # restarts, so persist it on a Longhorn volume.
  persistence:
    enabled: true
    size: 5Gi
    storageClassName: longhorn
  service:
    port: 80
  ingress:
    enabled: false

Manifests & templates

templates/grafana-datasources.yaml

Grafana datasource provisioning (Prometheus, Loki).

Show manifest
apiVersion: v1
kind: ConfigMap
metadata:
  name: {{ include "app.fullname" . }}-grafana-datasources
  namespace: {{ .Release.Namespace }}
  labels:
    {{- include "app.labels" . | nindent 4 }}
    # Picked up by the Grafana chart's datasource sidecar.
    grafana_datasource: "1"
data:
  datasources.yaml: |
    apiVersion: 1
    datasources:
      - name: Prometheus
        uid: prometheus
        type: prometheus
        access: proxy
        url: http://{{ include "app.fullname" . }}-prometheus-server.{{ .Release.Namespace }}.svc.cluster.local:9090
        isDefault: true
      - name: Loki
        uid: loki
        type: loki
        access: proxy
        url: http://{{ include "app.fullname" . }}-loki.{{ .Release.Namespace }}.svc.cluster.local:3100

templates/grafana-external-secret.yaml

ExternalSecret for Grafana admin and OAuth credentials.

Show manifest
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
  name: {{ include "app.fullname" . }}-grafana-admin
  namespace: {{ .Release.Namespace }}
  labels:
    {{- include "app.labels" . | nindent 4 }}
    app.kubernetes.io/component: secrets
spec:
  refreshInterval: 1h
  secretStoreRef:
    kind: ClusterSecretStore
    name: vault-cluster-secret-store
  target:
    name: monitoring-grafana-admin
    creationPolicy: Owner
  data:
    - secretKey: admin_email
      remoteRef:
        key: grafana
        property: admin_email
    - secretKey: admin_password
      remoteRef:
        key: grafana
        property: admin_password
    - secretKey: admin_user
      remoteRef:
        key: grafana
        property: admin_user
    - secretKey: secret_key
      remoteRef:
        key: grafana
        property: secret_key

templates/grafana-httproute.yaml

HTTPRoute for the Grafana UI.

Show manifest
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: {{ include "app.fullname" . }}-grafana-httproute
  namespace: {{ .Release.Namespace }}
  annotations:
    link.argocd.argoproj.io/external-link: "https://grafana.opi5cluster.co.uk"
spec:
  parentRefs:
    - name: istio-gateway
      namespace: istio
      sectionName: websecure
  hostnames:
    - grafana.opi5cluster.co.uk
  rules:
    - backendRefs:
        - name: {{ include "app.fullname" . }}-grafana
          port: 80
  • Loki and Alloy deploy as part of the same umbrella stack entry.

Trade-offs

Decision. 30-day retention on local storage

Alternative. A remote long-term metrics store (Thanos or Mimir)

Why. Lab-scale dashboards do not justify a second storage system; manifests are covered by Velero anyway.

← Back to Observability & Dashboard · All service groups