Whisper

Turns audio into text with local speech-to-text models

Whisper transcribes audio to text on demand, used for generating subtitles and searchable metadata.

Manifests & templates

templates/whisper/whisper-rollout.yaml

Whisper Rollout

Show manifest
{{- $projectName  := .Values.global.projectName -}}
{{- $hostMediaMountPath  := .Values.global.hostMediaMountPath -}}
{{- $timeZone := .Values.global.timeZone -}}
{{- $userID := .Values.global.userID -}}
{{- with .Values.applications.whisper }}
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: {{ .name  }}
  namespace: {{ $projectName }}
spec:
  replicas: 1
  strategy:
    canary:
      maxSurge: 0
      maxUnavailable: 1
      steps:
      - setWeight: 100
  selector:
    matchLabels:
      app: {{ .name  }}
  template:
    metadata:
      labels:
        app: {{ .name  }}
    spec:
      {{- include "media-systems.nodePlacement" $ | nindent 6 }}
      serviceAccount: {{ $projectName }}-service-account
      serviceAccountName: {{ $projectName }}-service-account
      automountServiceAccountToken: true
      securityContext:
        runAsUser: {{ $userID }}
        runAsGroup: {{ $userID }}
        fsGroup: {{ $userID }}
      containers:
      - name: {{ .name  }}
        image: {{ .image }}
        imagePullPolicy: IfNotPresent
        ports:
        - name: {{ .name | substr 0 9 }}-http
          containerPort: {{ .ports.http }}
          protocol: TCP
        env:
        - name: ASR_ENGINE
          value: faster_whisper
        - name: ASR_MODEL
          value: {{ .model | quote }}
        - name: ASR_DEVICE
          value: cpu
        - name: ASR_MODEL_PATH
          value: /cache/models
        - name: HF_HOME
          value: /cache
        - name: HOME
          value: /cache
        - name: TZ
          value: {{ $timeZone | quote }}
        # Silence Python 3.12 invalid-escape-sequence SyntaxWarnings from the
        # image's pyannote.database dependency (cosmetic; app unaffected).
        - name: PYTHONWARNINGS
          value: "ignore::SyntaxWarning"
        volumeMounts:
        - name: {{ .name }}-cache
          mountPath: /cache
        resources:
          requests:
            cpu: "2"
            memory: "2Gi"
      dnsPolicy: ClusterFirst
      restartPolicy: Always
      schedulerName: default-scheduler
      volumes:
      - name: {{ .name }}-cache
        hostPath:
          path: {{ $hostMediaMountPath }}/whisper-models
          type: Directory
{{- end -}}
  • media-systems is one umbrella Application: the whole pipeline deploys as a unit, pinned to a single node with local SSD storage. Per-app settings live in the media chart, not here.
  • This app is a section of the shared media-systems values file above; the media stack deploys as one ArgoCD Application.

Trade-offs

Decision. Local Whisper on the media node

Alternative. A cloud transcription API

Why. Audio never leaves the LAN, and the media node is otherwise idle between transcode passes.

← Back to Media · All service groups