Loki + Alloy

Collects logs from every container with Alloy and keeps them searchable for 14 days

Alloy collects container logs and ships them to Loki, which indexes only labels and keeps 14 days of log chunks. Queries run from Grafana next to the metrics.

Chart values

The full monitoring/values.yaml from the service's own chart:

# ─────────────────────────────────────────────────────────────────────
# Centralized monitoring: Grafana Alloy (DaemonSet) → Prometheus + Loki,
# dashboards in Grafana. Self-hosted; no external SaaS, no CRDs.
# ─────────────────────────────────────────────────────────────────────

# ── Cluster-wide metadata attached to every metric / log ─────────────
cluster:
  name: opi5-cluster
  environment: prod

# ── Static scrape targets for the cluster services ───────────────────
# Addresses are in-cluster Service DNS names. Each entry was verified
# against a rendered upstream chart (not guessed). If a service name
# ever changes, fix it here.
#
# Live verification:   kubectl get svc -A
#
# Some targets are only scraped once the corresponding chart change has
# been applied (see the per-chart "monitoring" commits):
#   * argo-cd / argo-workflows / argo-rollouts — metrics were off by
#     default and have been enabled in their values.
#   * vault — telemetry stanza added (path /v1/sys/metrics).
#   * openvino — metrics enabled on the REST port (`--metrics_enable`),
#     scraped from the Service's :8000/metrics.
metricsScrapeTargets:
  - name: istio-gateway
    job: istio-gateway
    url: istio-gateway-istio.istio.svc.cluster.local:15090
    namespace: istio
    component: gateway
  - name: argocd_application_controller
    job: argocd/application-controller
    url: argocd-application-controller-metrics.argocd.svc.cluster.local:8082
    namespace: argocd
    component: controller
  - name: argocd_server
    job: argocd/server
    url: argocd-server-metrics.argocd.svc.cluster.local:8083
    namespace: argocd
    component: server
  - name: argocd_repo_server
    job: argocd/repo-server
    url: argocd-repo-server-metrics.argocd.svc.cluster.local:8084
    namespace: argocd
    component: repo-server
  - name: argocd_applicationset_controller
    job: argocd/applicationset-controller
    url: argocd-applicationset-controller-metrics.argocd.svc.cluster.local:8080
    namespace: argocd
    component: applicationset-controller
  - name: argo_rollouts_controller
    job: argo-rollouts/controller
    url: argo-rollouts-metrics.argo-rollouts.svc.cluster.local:8090
    namespace: argo-rollouts
    component: controller
  - name: argo_workflows_controller
    job: argo-workflows/controller
    url: argo-workflows-controller.argo-workflows.svc.cluster.local:8080
    namespace: argo-workflows
    component: controller
  - name: argo_workflows_server
    job: argo-workflows/server
    url: argo-workflows-server.argo-workflows.svc.cluster.local:2746
    namespace: argo-workflows
    component: server
  - name: vault
    job: vault
    url: vault.vault.svc.cluster.local:8200
    namespace: vault
    component: server
    path: /v1/sys/metrics?format=prometheus
  - name: longhorn_manager
    job: longhorn/manager
    url: longhorn-backend.longhorn-system.svc.cluster.local:9500
    namespace: longhorn-system
    component: manager
  - name: openvino_model_server
    job: openvino/model-server
    url: openvino-service.openvino.svc.cluster.local:8000
    namespace: openvino
    component: model-server
    path: /metrics
  - name: zot
    job: zot/registry
    url: zot.zot.svc.cluster.local:5000
    namespace: zot
    component: registry
    path: /metrics

# ── Grafana Alloy (upstream chart) ───────────────────────────────────
# Note: the upstream alloy chart reads its runtime settings from
# `.Values.alloy` (and root-level `crds`/`controller`). Because Helm maps
# the parent's `alloy:` block to the subchart root, runtime settings need
# the nested `alloy.alloy:` key.
alloy:
  alloy:
    # We render our own ConfigMap (templates/alloy-config.yaml) so the
    # River config is generated from the values above. The ConfigMap name
    # defaults to the chart-computed fullname (`{release}-alloy`), so it
    # matches whatever the DaemonSet expects regardless of release name.
    configMap:
      create: false
    mounts:
      # Logs are streamed via the Kubernetes API (loki.source.kubernetes),
      # not tailed from the node filesystem — no host mounts required.
      varlog: false
      dockercontainers: false
    storagePath: /tmp/alloy
    enableReporting: false
    # Cluster mode: with one Alloy pod per node, every pod was scraping the
    # full cluster-wide target set, so Prometheus rejected all but the first
    # copy of each series ("out of order sample from remote write").
    # Clustering hash-shards the cluster-wide scrape targets so exactly one
    # member owns each target; node-local scrapes (kubelet/cAdvisor and
    # node-exporter pods) are instead scoped to the local node in the
    # Alloy config (see templates/alloy-config.yaml), where each node has
    # exactly one writer by construction.
    clustering:
      enabled: true
    resources:
      requests:
        cpu: 100m
        memory: 128Mi
      limits:
        cpu: 500m
        memory: 512Mi
    securityContext:
      allowPrivilegeEscalation: false
      capabilities:
        drop:
          - ALL
      seccompProfile:
        type: RuntimeDefault

  # The chart's default `crds` subchart installs the
  # monitoring.grafana.com/podlogs CRD — we don't use podlogs, so keep
  # the cluster CRD-free.
  crds:
    create: false

  # Alloy controller: DaemonSet, one pod per node.
  controller:
    type: daemonset
    # Run on every worker node, including the tainted data/LLM nodes.
    tolerations:
      - key: opi5.cluster/role
        operator: Equal
        value: data
        effect: NoSchedule
    nodeSelector: {}

# ── kube-state-metrics (upstream chart) ──────────────────────────────
kube-state-metrics:
  resources:
    requests:
      cpu: 20m
      memory: 80Mi
    limits:
      cpu: 100m
      memory: 200Mi

# ── Prometheus (upstream chart) ──────────────────────────────────────
# Pure metrics store + rule engine. Alloy is the only scraper and pushes
# via remote_write; Prometheus itself has no scrape configs to maintain.
# Only the server workload is deployed — Alertmanager, node-exporter,
# pushgateway and the bundled kube-state-metrics are all disabled.
prometheus:
  alertmanager:
    enabled: false
  kube-state-metrics:
    enabled: false
  # node-exporter provides node filesystem/network/load metrics for the
  # Host Overview dashboard. Alloy discovers its pods via the
  # prometheus.io/scrape annotations below (annotation-based discovery),
  # so no extra scrape config is needed.
  prometheus-node-exporter:
    enabled: true
    podAnnotations:
      prometheus.io/scrape: "true"
      prometheus.io/port: "19100"
    # hostNetwork pods bind the node's port directly. Traefik's metrics
    # entrypoint holds host 9100 (traefik chart default) on opi5-worker-2,
    # which kept this node's node-exporter Pending — so listen on 19100
    # instead. service.port drives both the bind address and containerPort.
    service:
      port: 19100
      targetPort: 19100
    # DaemonSet must cover every worker node for Host Overview. The broad
    # operator:Exists toleration is the chart default (any NoSchedule
    # taint); the explicit data-role toleration below documents the
    # intent to also run on the tainted data nodes.
    tolerations:
      - effect: NoSchedule
        operator: Exists
      - key: opi5.cluster/role
        operator: Equal
        value: data
        effect: NoSchedule
  prometheus-pushgateway:
    enabled: false
  server:
    retention: 30d
    # Accept remote_write from Alloy. Prometheus v3.13+ reverted to the
    # --web.enable-remote-write-receiver flag (the remote-write-receiver
    # entry was removed from --enable-feature), and the bundled
    # config-reloader needs --web.enable-lifecycle to trigger reloads
    # (403 otherwise).
    extraFlags:
      - web.enable-remote-write-receiver
      - web.enable-lifecycle
    service:
      servicePort: 9090
    persistentVolume:
      size: 20Gi

# ── Loki (upstream chart) ────────────────────────────────────────────
# Single-binary mode (SSD/SimpleScalable is deprecated and removed in
# Loki 4). Filesystem storage on a Longhorn PVC — no object store needed.
# Note: `loki:` here maps to the subchart root; the Loki config sections
# are under `loki.loki`.
loki:
  deploymentMode: SingleBinary
  singleBinary:
    replicas: 1
    persistence:
      size: 20Gi
  # Zero out the SimpleScalable targets (defaults are non-zero).
  write:
    replicas: 0
  read:
    replicas: 0
  backend:
    replicas: 0
  # Disable optional components we don't need.
  gateway:
    enabled: false
  lokiCanary:
    enabled: false
  test:
    enabled: false
  monitoring:
    selfMonitoring:
      enabled: false
    lokiCanary:
      enabled: false
  chunksCache:
    enabled: false
  resultsCache:
    enabled: false
  loki:
    # Single-tenant cluster: the chart default (auth_enabled: true) makes
    # Loki 401 every request without an X-Scope-OrgID header, which breaks
    # both Alloy's log pushes and Grafana's queries.
    auth_enabled: false
    commonConfig:
      # Single-binary single-replica: the chart default replication_factor
      # of 3 leaves the ring with "too many unhealthy instances" and every
      # query fails with 500.
      replication_factor: 1
    storage:
      type: filesystem
    schemaConfig:
      configs:
        - from: "2024-01-01"
          store: tsdb
          object_store: filesystem
          schema: v13
          index:
            prefix: index_
            period: 24h
    # 14-day log retention. Loki 3.x requires delete_request_store when
    # retention is enabled — "filesystem" names the local store.
    limits_config:
      retention_period: 336h
      # Enable the log-volume endpoint so Grafana Explore can render the
      # volume-over-time histogram.
      volume_enabled: true
    compactor:
      retention_enabled: true
      delete_request_store: filesystem

# ── Grafana (upstream chart) ─────────────────────────────────────────
# Admin credentials and the security secret_key come from Vault via ESO
# (Vault kv path `grafana` with properties admin_email, admin_password,
# admin_user, secret_key) — synced into the `monitoring-grafana-admin`
# secret by templates/grafana-external-secret.yaml. Keep the secret name
# in sync with that template's target.name.
grafana:
  admin:
    existingSecret: monitoring-grafana-admin
    userKey: admin_user
    passwordKey: admin_password
  # [security] secret_key — Grafana reads GF_SECURITY_SECRET_KEY from env.
  envValueFrom:
    GF_SECURITY_SECRET_KEY:
      secretKeyRef:
        name: monitoring-grafana-admin
        key: secret_key
  # The chart default readinessProbe has no initialDelaySeconds, so it
  # flaps "connection refused" while Grafana binds :3000 on every start
  # (slower on a fresh Longhorn PVC + first-boot migrations).
  readinessProbe:
    initialDelaySeconds: 20
  sidecar:
    datasources:
      enabled: true
      label: grafana_datasource
    # Load dashboards-as-code from ConfigMaps labelled
    # `grafana_dashboard: "1"` (see templates/grafana-dashboards/).
    dashboards:
      enabled: true
      label: grafana_dashboard
      # Absolute path (the chart default is /tmp/dashboards). A relative
      # value makes the sidecar resolve it against /app (its WORKDIR) and
      # fail to write, crash-looping grafana-sc-dashboard.
      folder: /Hosts
  # Dashboards/datasources are provisioned via ConfigMap, but the SQLite
  # DB (annotations, alert rules, orgs, folders) must survive pod
  # restarts, so persist it on a Longhorn volume.
  persistence:
    enabled: true
    size: 5Gi
    storageClassName: longhorn
  service:
    port: 80
  ingress:
    enabled: false

Manifests & templates

templates/alloy-config.yaml

Alloy pipeline config: log discovery across nodes, processing, and shipping into Loki with 14-day retention.

Show manifest
apiVersion: v1
kind: ConfigMap
metadata:
  name: {{ include "app.fullname" . }}-alloy
  namespace: {{ .Release.Namespace }}
  labels:
    {{- include "app.labels" . | nindent 4 }}
    app.kubernetes.io/component: collector
data:
  config.alloy: |-
{{- $cluster := .Values.cluster }}
{{- $targets := .Values.metricsScrapeTargets }}
{{- $ksm := printf "%s-kube-state-metrics.%s.svc.cluster.local:8080" (include "app.fullname" .) .Release.Namespace }}
{{- $prom := printf "http://%s-prometheus-server.%s.svc.cluster.local:9090/api/v1/write" (include "app.fullname" .) .Release.Namespace }}
{{- $loki := printf "http://%s-loki.%s.svc.cluster.local:3100/loki/api/v1/push" (include "app.fullname" .) .Release.Namespace }}
    // Grafana Alloy — collection layer → self-hosted stack
    //
    // Applications never learn about the backend: they write JSON logs to
    // stdout and (optionally) expose /metrics behind prometheus.io
    // annotations. Alloy is the single collector:
    //   metrics → Prometheus (remote_write)
    //   logs    → Loki (loki.write)

    // ── Discovery ───────────────────────────────────────────────────────

    discovery.kubernetes "nodes" {
      role = "node"
    }

    discovery.kubernetes "pods" {
      role = "pod"
    }

    // ── Egress: metrics → in-cluster Prometheus ─────────────────────────

    prometheus.remote_write "prometheus" {
      external_labels = {
        cluster     = "{{ $cluster.name }}",
        environment = "{{ $cluster.environment }}",
      }

      endpoint {
        url            = "{{ $prom }}"
        remote_timeout = "30s"
      }
    }

    // ── Egress: logs → in-cluster Loki ──────────────────────────────────

    loki.write "local" {
      endpoint {
        url = "{{ $loki }}"
      }
    }

    // ── Metrics: kubelet + cAdvisor (node / container health) ──────────

    discovery.relabel "node_targets" {
      targets = discovery.kubernetes.nodes.targets

      // Scope to the local node only. Without this filter every Alloy pod
      // scrapes every node's kubelet, and N Alloys writing the same series
      // with identical timestamps makes Prometheus reject all but the
      // first copy ("out of order sample from remote write").
      rule {
        source_labels = ["__meta_kubernetes_node_name"]
        regex         = sys.env("K8S_NODE_NAME")
        action        = "keep"
      }
      // `discovery.kubernetes` role=node sets __address__ to the node IP
      // without a port; append the kubelet HTTPS metrics port.
      rule {
        source_labels = ["__address__"]
        regex         = "([^:]+)(?::\\d+)?"
        replacement   = "$1:10250"
        target_label  = "__address__"
      }
      rule {
        source_labels = ["__meta_kubernetes_node_name"]
        target_label  = "instance"
      }
    }

    prometheus.scrape "kubelet" {
      job_name = "integrations/kubernetes/kubelet"
      targets = discovery.relabel.node_targets.output
      scheme            = "https"
      bearer_token_file = "/var/run/secrets/kubernetes.io/serviceaccount/token"
      tls_config {
        ca_file              = "/var/run/secrets/kubernetes.io/serviceaccount/ca.crt"
        insecure_skip_verify = true
      }
      metrics_path   = "/metrics"
      scrape_interval = "30s"
      forward_to     = [prometheus.remote_write.prometheus.receiver]
    }

    prometheus.scrape "cadvisor" {
      job_name = "integrations/kubernetes/cadvisor"
      targets = discovery.relabel.node_targets.output
      scheme            = "https"
      bearer_token_file = "/var/run/secrets/kubernetes.io/serviceaccount/token"
      tls_config {
        ca_file              = "/var/run/secrets/kubernetes.io/serviceaccount/ca.crt"
        insecure_skip_verify = true
      }
      metrics_path   = "/metrics/cadvisor"
      scrape_interval = "30s"
      forward_to     = [prometheus.remote_write.prometheus.receiver]
    }

    // ── Metrics: kube-state-metrics (cluster object state) ──────────────

    prometheus.scrape "kube_state_metrics" {
      job_name = "kube-state-metrics"
      targets = [{
        __address__ = "{{ $ksm }}",
        namespace   = "{{ .Release.Namespace }}",
        component   = "kube-state-metrics",
      }]
      scrape_interval = "30s"
      // Cluster-wide target: identical on every member, so cluster mode
      // assigns it to exactly one Alloy (no duplicate series).
      clustering {
        enabled = true
      }
      forward_to      = [prometheus.remote_write.prometheus.receiver]
    }

{{- range $t := $targets }}

    // ── Metrics: {{ $t.job }} ─────────────────────────────────────────────

    prometheus.scrape "{{ $t.name }}" {
      job_name = "{{ $t.job }}"
      targets = [{
        __address__ = "{{ $t.url }}",
        namespace   = "{{ $t.namespace }}",
        component   = "{{ $t.component | default $t.name }}",
      }]
{{- if $t.path }}
      metrics_path = "{{ $t.path }}"
{{- end }}
      scrape_interval = "30s"
      // Cluster-wide target: identical on every member, so cluster mode
      // assigns it to exactly one Alloy (no duplicate series).
      clustering {
        enabled = true
      }
      forward_to      = [prometheus.remote_write.prometheus.receiver]
    }
{{- end }}

    // ── Metrics: annotated pods (future bot fleet) ──────────────────────

    discovery.relabel "pod_metrics" {
      targets = discovery.kubernetes.pods.targets

      // Scope to the local node only — same rationale as node_targets:
      // without this, every Alloy pod scrapes every annotated pod on the
      // cluster and Prometheus rejects the duplicate series.
      rule {
        source_labels = ["__meta_kubernetes_pod_node_name"]
        regex         = sys.env("K8S_NODE_NAME")
        action        = "keep"
      }
      rule {
        source_labels = ["__meta_kubernetes_pod_annotation_prometheus_io_scrape"]
        action        = "keep"
        regex         = "true"
      }
      rule {
        source_labels = ["__address__", "__meta_kubernetes_pod_annotation_prometheus_io_port"]
        regex         = "([^:]+)(?::\\d+)?;(\\d+)"
        replacement   = "$1:$2"
        target_label  = "__address__"
      }
      rule {
        source_labels = ["__meta_kubernetes_pod_annotation_prometheus_io_path"]
        regex         = "(.+)"
        action        = "replace"
        target_label  = "__metrics_path__"
      }
      rule {
        source_labels = ["__meta_kubernetes_pod_name"]
        target_label  = "pod"
      }
      rule {
        source_labels = ["__meta_kubernetes_namespace"]
        target_label  = "namespace"
      }
      rule {
        source_labels = ["__meta_kubernetes_pod_container_name"]
        target_label  = "container"
      }
      rule {
        source_labels = ["__meta_kubernetes_pod_label_app"]
        target_label  = "app"
      }
      rule {
        source_labels = ["__meta_kubernetes_pod_node_name"]
        target_label  = "node"
      }
    }

    prometheus.scrape "pods" {
      job_name        = "integrations/kubernetes/pods"
      targets         = discovery.relabel.pod_metrics.output
      scrape_interval = "30s"
      forward_to      = [prometheus.remote_write.prometheus.receiver]
    }

    // ── Logs: local pods → Loki ─────────────────────────────────────────
    //
    // Pod logs are streamed over the Kubernetes API (loki.source.kubernetes)
    // rather than tailed from /var/log/pods. This cluster runs K3S with a
    // relocated data-dir (/mnt/ssd-small/rancher/k3s), so /var/log/pods is a
    // symlink into a path the DaemonSet doesn't mount — file-based tailing
    // fails on every pod. The API approach is layout-independent.
    //
    // Each Alloy instance must only tail pods scheduled on ITS OWN node.
    // Without this filter every node's Alloy follows every pod in the
    // cluster, and each follow session pins one inotify instance in the
    // kubelet. That fan-out (#alloys × #containers) exhausts
    // fs.inotify.max_user_instances on the node and breaks unrelated
    // inotify users (EMFILE).

    discovery.relabel "log_metadata" {
      targets = discovery.kubernetes.pods.targets

      rule {
        source_labels = ["__meta_kubernetes_pod_node_name"]
        regex         = sys.env("K8S_NODE_NAME")
        action        = "keep"
      }
      rule {
        source_labels = ["__meta_kubernetes_pod_container_name"]
        target_label  = "container"
      }
      rule {
        source_labels = ["__meta_kubernetes_pod_name"]
        target_label  = "pod"
      }
      rule {
        source_labels = ["__meta_kubernetes_namespace"]
        target_label  = "namespace"
      }
      rule {
        source_labels = ["__meta_kubernetes_pod_label_app"]
        target_label  = "app"
      }
    }

    loki.source.kubernetes "pod_logs" {
      targets    = discovery.relabel.log_metadata.output
      forward_to = [loki.process.pod_logs.receiver]
    }

    loki.process "pod_logs" {
      // Parse single-line JSON logs (Traefik, Longhorn, Muse, Whistle …)
      // into fields. Plain-text lines pass through unchanged.
      stage.json {
        expressions = {
          level  = "level",
          msg    = "msg",
          logger = "logger",
        }
      }
      // Only the bounded `level` field becomes a label — everything else
      // stays in the log body to protect Loki cardinality.
      stage.labels {
        values = {
          level = "level",
        }
      }
      forward_to = [loki.write.local.receiver]
    }

Trade-offs

Decision. Loki

Alternative. Elasticsearch or OpenSearch

Why. Index-free storage is dramatically cheaper on arm64 nodes with small disks, and label-based queries cover everything the lab needs.

← Back to Observability & Dashboard · All service groups