Decision. Loki
Alternative. Elasticsearch or OpenSearch
Why. Index-free storage is dramatically cheaper on arm64 nodes with small disks, and label-based queries cover everything the lab needs.
Collects logs from every container with Alloy and keeps them searchable for 14 days
Alloy collects container logs and ships them to Loki, which indexes only labels and keeps 14 days of log chunks. Queries run from Grafana next to the metrics.
The full monitoring/values.yaml from the service's
own chart:
# ─────────────────────────────────────────────────────────────────────
# Centralized monitoring: Grafana Alloy (DaemonSet) → Prometheus + Loki,
# dashboards in Grafana. Self-hosted; no external SaaS, no CRDs.
# ─────────────────────────────────────────────────────────────────────
# ── Cluster-wide metadata attached to every metric / log ─────────────
cluster:
name: opi5-cluster
environment: prod
# ── Static scrape targets for the cluster services ───────────────────
# Addresses are in-cluster Service DNS names. Each entry was verified
# against a rendered upstream chart (not guessed). If a service name
# ever changes, fix it here.
#
# Live verification: kubectl get svc -A
#
# Some targets are only scraped once the corresponding chart change has
# been applied (see the per-chart "monitoring" commits):
# * argo-cd / argo-workflows / argo-rollouts — metrics were off by
# default and have been enabled in their values.
# * vault — telemetry stanza added (path /v1/sys/metrics).
# * openvino — metrics enabled on the REST port (`--metrics_enable`),
# scraped from the Service's :8000/metrics.
metricsScrapeTargets:
- name: istio-gateway
job: istio-gateway
url: istio-gateway-istio.istio.svc.cluster.local:15090
namespace: istio
component: gateway
- name: argocd_application_controller
job: argocd/application-controller
url: argocd-application-controller-metrics.argocd.svc.cluster.local:8082
namespace: argocd
component: controller
- name: argocd_server
job: argocd/server
url: argocd-server-metrics.argocd.svc.cluster.local:8083
namespace: argocd
component: server
- name: argocd_repo_server
job: argocd/repo-server
url: argocd-repo-server-metrics.argocd.svc.cluster.local:8084
namespace: argocd
component: repo-server
- name: argocd_applicationset_controller
job: argocd/applicationset-controller
url: argocd-applicationset-controller-metrics.argocd.svc.cluster.local:8080
namespace: argocd
component: applicationset-controller
- name: argo_rollouts_controller
job: argo-rollouts/controller
url: argo-rollouts-metrics.argo-rollouts.svc.cluster.local:8090
namespace: argo-rollouts
component: controller
- name: argo_workflows_controller
job: argo-workflows/controller
url: argo-workflows-controller.argo-workflows.svc.cluster.local:8080
namespace: argo-workflows
component: controller
- name: argo_workflows_server
job: argo-workflows/server
url: argo-workflows-server.argo-workflows.svc.cluster.local:2746
namespace: argo-workflows
component: server
- name: vault
job: vault
url: vault.vault.svc.cluster.local:8200
namespace: vault
component: server
path: /v1/sys/metrics?format=prometheus
- name: longhorn_manager
job: longhorn/manager
url: longhorn-backend.longhorn-system.svc.cluster.local:9500
namespace: longhorn-system
component: manager
- name: openvino_model_server
job: openvino/model-server
url: openvino-service.openvino.svc.cluster.local:8000
namespace: openvino
component: model-server
path: /metrics
- name: zot
job: zot/registry
url: zot.zot.svc.cluster.local:5000
namespace: zot
component: registry
path: /metrics
# ── Grafana Alloy (upstream chart) ───────────────────────────────────
# Note: the upstream alloy chart reads its runtime settings from
# `.Values.alloy` (and root-level `crds`/`controller`). Because Helm maps
# the parent's `alloy:` block to the subchart root, runtime settings need
# the nested `alloy.alloy:` key.
alloy:
alloy:
# We render our own ConfigMap (templates/alloy-config.yaml) so the
# River config is generated from the values above. The ConfigMap name
# defaults to the chart-computed fullname (`{release}-alloy`), so it
# matches whatever the DaemonSet expects regardless of release name.
configMap:
create: false
mounts:
# Logs are streamed via the Kubernetes API (loki.source.kubernetes),
# not tailed from the node filesystem — no host mounts required.
varlog: false
dockercontainers: false
storagePath: /tmp/alloy
enableReporting: false
# Cluster mode: with one Alloy pod per node, every pod was scraping the
# full cluster-wide target set, so Prometheus rejected all but the first
# copy of each series ("out of order sample from remote write").
# Clustering hash-shards the cluster-wide scrape targets so exactly one
# member owns each target; node-local scrapes (kubelet/cAdvisor and
# node-exporter pods) are instead scoped to the local node in the
# Alloy config (see templates/alloy-config.yaml), where each node has
# exactly one writer by construction.
clustering:
enabled: true
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 512Mi
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop:
- ALL
seccompProfile:
type: RuntimeDefault
# The chart's default `crds` subchart installs the
# monitoring.grafana.com/podlogs CRD — we don't use podlogs, so keep
# the cluster CRD-free.
crds:
create: false
# Alloy controller: DaemonSet, one pod per node.
controller:
type: daemonset
# Run on every worker node, including the tainted data/LLM nodes.
tolerations:
- key: opi5.cluster/role
operator: Equal
value: data
effect: NoSchedule
nodeSelector: {}
# ── kube-state-metrics (upstream chart) ──────────────────────────────
kube-state-metrics:
resources:
requests:
cpu: 20m
memory: 80Mi
limits:
cpu: 100m
memory: 200Mi
# ── Prometheus (upstream chart) ──────────────────────────────────────
# Pure metrics store + rule engine. Alloy is the only scraper and pushes
# via remote_write; Prometheus itself has no scrape configs to maintain.
# Only the server workload is deployed — Alertmanager, node-exporter,
# pushgateway and the bundled kube-state-metrics are all disabled.
prometheus:
alertmanager:
enabled: false
kube-state-metrics:
enabled: false
# node-exporter provides node filesystem/network/load metrics for the
# Host Overview dashboard. Alloy discovers its pods via the
# prometheus.io/scrape annotations below (annotation-based discovery),
# so no extra scrape config is needed.
prometheus-node-exporter:
enabled: true
podAnnotations:
prometheus.io/scrape: "true"
prometheus.io/port: "19100"
# hostNetwork pods bind the node's port directly. Traefik's metrics
# entrypoint holds host 9100 (traefik chart default) on opi5-worker-2,
# which kept this node's node-exporter Pending — so listen on 19100
# instead. service.port drives both the bind address and containerPort.
service:
port: 19100
targetPort: 19100
# DaemonSet must cover every worker node for Host Overview. The broad
# operator:Exists toleration is the chart default (any NoSchedule
# taint); the explicit data-role toleration below documents the
# intent to also run on the tainted data nodes.
tolerations:
- effect: NoSchedule
operator: Exists
- key: opi5.cluster/role
operator: Equal
value: data
effect: NoSchedule
prometheus-pushgateway:
enabled: false
server:
retention: 30d
# Accept remote_write from Alloy. Prometheus v3.13+ reverted to the
# --web.enable-remote-write-receiver flag (the remote-write-receiver
# entry was removed from --enable-feature), and the bundled
# config-reloader needs --web.enable-lifecycle to trigger reloads
# (403 otherwise).
extraFlags:
- web.enable-remote-write-receiver
- web.enable-lifecycle
service:
servicePort: 9090
persistentVolume:
size: 20Gi
# ── Loki (upstream chart) ────────────────────────────────────────────
# Single-binary mode (SSD/SimpleScalable is deprecated and removed in
# Loki 4). Filesystem storage on a Longhorn PVC — no object store needed.
# Note: `loki:` here maps to the subchart root; the Loki config sections
# are under `loki.loki`.
loki:
deploymentMode: SingleBinary
singleBinary:
replicas: 1
persistence:
size: 20Gi
# Zero out the SimpleScalable targets (defaults are non-zero).
write:
replicas: 0
read:
replicas: 0
backend:
replicas: 0
# Disable optional components we don't need.
gateway:
enabled: false
lokiCanary:
enabled: false
test:
enabled: false
monitoring:
selfMonitoring:
enabled: false
lokiCanary:
enabled: false
chunksCache:
enabled: false
resultsCache:
enabled: false
loki:
# Single-tenant cluster: the chart default (auth_enabled: true) makes
# Loki 401 every request without an X-Scope-OrgID header, which breaks
# both Alloy's log pushes and Grafana's queries.
auth_enabled: false
commonConfig:
# Single-binary single-replica: the chart default replication_factor
# of 3 leaves the ring with "too many unhealthy instances" and every
# query fails with 500.
replication_factor: 1
storage:
type: filesystem
schemaConfig:
configs:
- from: "2024-01-01"
store: tsdb
object_store: filesystem
schema: v13
index:
prefix: index_
period: 24h
# 14-day log retention. Loki 3.x requires delete_request_store when
# retention is enabled — "filesystem" names the local store.
limits_config:
retention_period: 336h
# Enable the log-volume endpoint so Grafana Explore can render the
# volume-over-time histogram.
volume_enabled: true
compactor:
retention_enabled: true
delete_request_store: filesystem
# ── Grafana (upstream chart) ─────────────────────────────────────────
# Admin credentials and the security secret_key come from Vault via ESO
# (Vault kv path `grafana` with properties admin_email, admin_password,
# admin_user, secret_key) — synced into the `monitoring-grafana-admin`
# secret by templates/grafana-external-secret.yaml. Keep the secret name
# in sync with that template's target.name.
grafana:
admin:
existingSecret: monitoring-grafana-admin
userKey: admin_user
passwordKey: admin_password
# [security] secret_key — Grafana reads GF_SECURITY_SECRET_KEY from env.
envValueFrom:
GF_SECURITY_SECRET_KEY:
secretKeyRef:
name: monitoring-grafana-admin
key: secret_key
# The chart default readinessProbe has no initialDelaySeconds, so it
# flaps "connection refused" while Grafana binds :3000 on every start
# (slower on a fresh Longhorn PVC + first-boot migrations).
readinessProbe:
initialDelaySeconds: 20
sidecar:
datasources:
enabled: true
label: grafana_datasource
# Load dashboards-as-code from ConfigMaps labelled
# `grafana_dashboard: "1"` (see templates/grafana-dashboards/).
dashboards:
enabled: true
label: grafana_dashboard
# Absolute path (the chart default is /tmp/dashboards). A relative
# value makes the sidecar resolve it against /app (its WORKDIR) and
# fail to write, crash-looping grafana-sc-dashboard.
folder: /Hosts
# Dashboards/datasources are provisioned via ConfigMap, but the SQLite
# DB (annotations, alert rules, orgs, folders) must survive pod
# restarts, so persist it on a Longhorn volume.
persistence:
enabled: true
size: 5Gi
storageClassName: longhorn
service:
port: 80
ingress:
enabled: false templates/alloy-config.yaml Alloy pipeline config: log discovery across nodes, processing, and shipping into Loki with 14-day retention.
apiVersion: v1
kind: ConfigMap
metadata:
name: {{ include "app.fullname" . }}-alloy
namespace: {{ .Release.Namespace }}
labels:
{{- include "app.labels" . | nindent 4 }}
app.kubernetes.io/component: collector
data:
config.alloy: |-
{{- $cluster := .Values.cluster }}
{{- $targets := .Values.metricsScrapeTargets }}
{{- $ksm := printf "%s-kube-state-metrics.%s.svc.cluster.local:8080" (include "app.fullname" .) .Release.Namespace }}
{{- $prom := printf "http://%s-prometheus-server.%s.svc.cluster.local:9090/api/v1/write" (include "app.fullname" .) .Release.Namespace }}
{{- $loki := printf "http://%s-loki.%s.svc.cluster.local:3100/loki/api/v1/push" (include "app.fullname" .) .Release.Namespace }}
// Grafana Alloy — collection layer → self-hosted stack
//
// Applications never learn about the backend: they write JSON logs to
// stdout and (optionally) expose /metrics behind prometheus.io
// annotations. Alloy is the single collector:
// metrics → Prometheus (remote_write)
// logs → Loki (loki.write)
// ── Discovery ───────────────────────────────────────────────────────
discovery.kubernetes "nodes" {
role = "node"
}
discovery.kubernetes "pods" {
role = "pod"
}
// ── Egress: metrics → in-cluster Prometheus ─────────────────────────
prometheus.remote_write "prometheus" {
external_labels = {
cluster = "{{ $cluster.name }}",
environment = "{{ $cluster.environment }}",
}
endpoint {
url = "{{ $prom }}"
remote_timeout = "30s"
}
}
// ── Egress: logs → in-cluster Loki ──────────────────────────────────
loki.write "local" {
endpoint {
url = "{{ $loki }}"
}
}
// ── Metrics: kubelet + cAdvisor (node / container health) ──────────
discovery.relabel "node_targets" {
targets = discovery.kubernetes.nodes.targets
// Scope to the local node only. Without this filter every Alloy pod
// scrapes every node's kubelet, and N Alloys writing the same series
// with identical timestamps makes Prometheus reject all but the
// first copy ("out of order sample from remote write").
rule {
source_labels = ["__meta_kubernetes_node_name"]
regex = sys.env("K8S_NODE_NAME")
action = "keep"
}
// `discovery.kubernetes` role=node sets __address__ to the node IP
// without a port; append the kubelet HTTPS metrics port.
rule {
source_labels = ["__address__"]
regex = "([^:]+)(?::\\d+)?"
replacement = "$1:10250"
target_label = "__address__"
}
rule {
source_labels = ["__meta_kubernetes_node_name"]
target_label = "instance"
}
}
prometheus.scrape "kubelet" {
job_name = "integrations/kubernetes/kubelet"
targets = discovery.relabel.node_targets.output
scheme = "https"
bearer_token_file = "/var/run/secrets/kubernetes.io/serviceaccount/token"
tls_config {
ca_file = "/var/run/secrets/kubernetes.io/serviceaccount/ca.crt"
insecure_skip_verify = true
}
metrics_path = "/metrics"
scrape_interval = "30s"
forward_to = [prometheus.remote_write.prometheus.receiver]
}
prometheus.scrape "cadvisor" {
job_name = "integrations/kubernetes/cadvisor"
targets = discovery.relabel.node_targets.output
scheme = "https"
bearer_token_file = "/var/run/secrets/kubernetes.io/serviceaccount/token"
tls_config {
ca_file = "/var/run/secrets/kubernetes.io/serviceaccount/ca.crt"
insecure_skip_verify = true
}
metrics_path = "/metrics/cadvisor"
scrape_interval = "30s"
forward_to = [prometheus.remote_write.prometheus.receiver]
}
// ── Metrics: kube-state-metrics (cluster object state) ──────────────
prometheus.scrape "kube_state_metrics" {
job_name = "kube-state-metrics"
targets = [{
__address__ = "{{ $ksm }}",
namespace = "{{ .Release.Namespace }}",
component = "kube-state-metrics",
}]
scrape_interval = "30s"
// Cluster-wide target: identical on every member, so cluster mode
// assigns it to exactly one Alloy (no duplicate series).
clustering {
enabled = true
}
forward_to = [prometheus.remote_write.prometheus.receiver]
}
{{- range $t := $targets }}
// ── Metrics: {{ $t.job }} ─────────────────────────────────────────────
prometheus.scrape "{{ $t.name }}" {
job_name = "{{ $t.job }}"
targets = [{
__address__ = "{{ $t.url }}",
namespace = "{{ $t.namespace }}",
component = "{{ $t.component | default $t.name }}",
}]
{{- if $t.path }}
metrics_path = "{{ $t.path }}"
{{- end }}
scrape_interval = "30s"
// Cluster-wide target: identical on every member, so cluster mode
// assigns it to exactly one Alloy (no duplicate series).
clustering {
enabled = true
}
forward_to = [prometheus.remote_write.prometheus.receiver]
}
{{- end }}
// ── Metrics: annotated pods (future bot fleet) ──────────────────────
discovery.relabel "pod_metrics" {
targets = discovery.kubernetes.pods.targets
// Scope to the local node only — same rationale as node_targets:
// without this, every Alloy pod scrapes every annotated pod on the
// cluster and Prometheus rejects the duplicate series.
rule {
source_labels = ["__meta_kubernetes_pod_node_name"]
regex = sys.env("K8S_NODE_NAME")
action = "keep"
}
rule {
source_labels = ["__meta_kubernetes_pod_annotation_prometheus_io_scrape"]
action = "keep"
regex = "true"
}
rule {
source_labels = ["__address__", "__meta_kubernetes_pod_annotation_prometheus_io_port"]
regex = "([^:]+)(?::\\d+)?;(\\d+)"
replacement = "$1:$2"
target_label = "__address__"
}
rule {
source_labels = ["__meta_kubernetes_pod_annotation_prometheus_io_path"]
regex = "(.+)"
action = "replace"
target_label = "__metrics_path__"
}
rule {
source_labels = ["__meta_kubernetes_pod_name"]
target_label = "pod"
}
rule {
source_labels = ["__meta_kubernetes_namespace"]
target_label = "namespace"
}
rule {
source_labels = ["__meta_kubernetes_pod_container_name"]
target_label = "container"
}
rule {
source_labels = ["__meta_kubernetes_pod_label_app"]
target_label = "app"
}
rule {
source_labels = ["__meta_kubernetes_pod_node_name"]
target_label = "node"
}
}
prometheus.scrape "pods" {
job_name = "integrations/kubernetes/pods"
targets = discovery.relabel.pod_metrics.output
scrape_interval = "30s"
forward_to = [prometheus.remote_write.prometheus.receiver]
}
// ── Logs: local pods → Loki ─────────────────────────────────────────
//
// Pod logs are streamed over the Kubernetes API (loki.source.kubernetes)
// rather than tailed from /var/log/pods. This cluster runs K3S with a
// relocated data-dir (/mnt/ssd-small/rancher/k3s), so /var/log/pods is a
// symlink into a path the DaemonSet doesn't mount — file-based tailing
// fails on every pod. The API approach is layout-independent.
//
// Each Alloy instance must only tail pods scheduled on ITS OWN node.
// Without this filter every node's Alloy follows every pod in the
// cluster, and each follow session pins one inotify instance in the
// kubelet. That fan-out (#alloys × #containers) exhausts
// fs.inotify.max_user_instances on the node and breaks unrelated
// inotify users (EMFILE).
discovery.relabel "log_metadata" {
targets = discovery.kubernetes.pods.targets
rule {
source_labels = ["__meta_kubernetes_pod_node_name"]
regex = sys.env("K8S_NODE_NAME")
action = "keep"
}
rule {
source_labels = ["__meta_kubernetes_pod_container_name"]
target_label = "container"
}
rule {
source_labels = ["__meta_kubernetes_pod_name"]
target_label = "pod"
}
rule {
source_labels = ["__meta_kubernetes_namespace"]
target_label = "namespace"
}
rule {
source_labels = ["__meta_kubernetes_pod_label_app"]
target_label = "app"
}
}
loki.source.kubernetes "pod_logs" {
targets = discovery.relabel.log_metadata.output
forward_to = [loki.process.pod_logs.receiver]
}
loki.process "pod_logs" {
// Parse single-line JSON logs (Traefik, Longhorn, Muse, Whistle …)
// into fields. Plain-text lines pass through unchanged.
stage.json {
expressions = {
level = "level",
msg = "msg",
logger = "logger",
}
}
// Only the bounded `level` field becomes a label — everything else
// stays in the log body to protect Loki cardinality.
stage.labels {
values = {
level = "level",
}
}
forward_to = [loki.write.local.receiver]
} Decision. Loki
Alternative. Elasticsearch or OpenSearch
Why. Index-free storage is dramatically cheaper on arm64 nodes with small disks, and label-based queries cover everything the lab needs.