OpenVINO Model Server

Self-hosted LLM, embedding, and image models behind an OpenAI-compatible API

OpenVINO Model Server hosts the local model stack: LLMs, embedding models, and image models behind a single OpenAI-compatible endpoint. Consumers switch models by name, not by deployment.

ArgoCD configuration

Excerpt from argocd-apps/values.yaml in the argocd-apps chart, with annotations added for this site:

openvino:
  application: true
  project: applications
  # false: model servers are memory-hungry; upgrades are scheduled
  autoSync: false

Chart values

The full openvino/values.yaml from the service's own chart:

runOnHost: opi5-worker-5-genai
image: registry.opi5cluster.co.uk/openvino:2026.3-gpu
url: openvino.opi5cluster.co.uk
port: 8000
# Monitoring: OVMS serves Prometheus metrics on the REST port (`--metrics_enable`),
# scraped by the `monitoring` Alloy DaemonSet → Prometheus/Grafana.
volume:
  storageClass: longhorn-ssd-large
  size: 300Gi
  accessModes: ReadWriteMany
resources:
  requests:
    cpu: 6
    memory: 10Gi
modelsToDownload:
  Qwen3.5-4B:
    modelRepo: OpenVINO/Qwen3.5-4B-int4-ov
    modelTask: text_generation
    targetDevice: GPU

  Qwen3.6-27B:
    modelRepo: OpenVINO/Qwen3.6-27B-int4-ov
    modelTask: text_generation
    targetDevice: GPU
    toolParser: qwen3coder
    reasoningParser: qwen3

  Qwen3-Embedding:
    modelRepo: OpenVINO/Qwen3-Embedding-0.6B-int8-ov
    modelTask: embeddings
    targetDevice: GPU

  Qwen3-Reranker:
    modelRepo: OpenVINO/Qwen3-Reranker-0.6B-int8-ov
    modelTask: rerank
    targetDevice: GPU

  FLUX:
    modelRepo: OpenVINO/FLUX.1-schnell-int4-ov
    modelTask: image_generation
    targetDevice: CPU
    defaultNumInferenceSteps: 4

tolerations:
  - key: opi5.cluster/role
    operator: Equal
    value: data
    effect: NoSchedule

env:
  - name: IPEX_DEVICES
    value: GPU
  - name: DPCPP_DEVICE_TYPE
    value: GPU

Manifests & templates

templates/config.yaml

OpenVINO Model Server config: served models and plugin settings.

Show manifest
apiVersion: v1
kind: ConfigMap
metadata:
  name: openvino-configmap
  namespace: {{ .Release.Namespace }}
data:
  config.json: |
    {
      "model_config_list": [
        {{- $models := .Values.modelsToDownload -}}
        {{- $first := true -}}
        {{- range $name, $value := $models }}
            {{- if not $first }},{{ end }}
            {
              "config": {
                "name": "openai/{{ $name }}",
                "base_path": "/models/{{ $value.modelRepo }}"
              }
            }
            {{- $first = false -}}
        {{- end }}
      ]
    }

templates/external-secret.yaml

ExternalSecret syncing the service credentials from Vault into the namespace.

Show manifest
# Create External Service to extract values from Doppler
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
  name: {{ .Release.Namespace }}-external-secret
  namespace: {{ .Release.Namespace }}
spec:
  refreshInterval: 30s
  secretStoreRef:
    kind: ClusterSecretStore
    name: vault-cluster-secret-store
  target:
    name: {{ .Release.Namespace }}-secret
    creationPolicy: Owner
  data:
    - secretKey: OPENVINO_API_KEY
      remoteRef:
        key: genai
        property: openvino-api-key
    - secretKey: HF_TOKEN
      remoteRef:
        key: huggingface
        property: token

templates/http-route.yaml

Gateway API HTTPRoute exposing the service through the Istio gateway under opi5cluster.co.uk.

Show manifest
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: {{ .Release.Namespace  }}-httproute
  namespace: {{ .Release.Namespace }}
  annotations:
    link.argocd.argoproj.io/external-link: "https://{{ .Values.url }}"
spec:
  parentRefs:
    - name: istio-gateway
      namespace: istio
      sectionName: websecure
  hostnames:
    - {{ .Values.url }}
  rules:
    - backendRefs:
        - name: {{ .Release.Namespace }}-service
          port: {{ .Values.port }}

templates/models-downloader-wf.yaml

Argo Workflow downloading models into the OpenVINO model storage.

Show manifest
{{- $nameSpace  := .Release.Namespace }}
{{- range $model, $property := .Values.modelsToDownload }}
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
  name: {{ $model | lower }}-model-downloader-wf
spec:
  entrypoint: openvino-model-downloader
  serviceAccountName: argo-workflow-sa
  arguments:
    parameters:
    - name: modelName
      value: "{{ $model }}"
    - name: modelRepo
      value: "{{ $property.modelRepo | default "" }}"
    - name: modelTask
      value: "{{ $property.modelTask | default "" }}"
    - name: targetDevice
      value: "{{ $property.targetDevice | default "" }}"
    - name: toolParser
      value: "{{ $property.toolParser | default "" }}"
    - name: reasoningParser
      value: "{{ $property.reasoningParser | default "" }}"
    - name: ggufFilename
      value: "{{ $property.ggufFilename | default "" }}"
    - name: defaultNumInferenceSteps
      value: "{{ $property.defaultNumInferenceSteps | default "" }}"
  templates:
    - name: openvino-model-downloader
      steps:
        - - name: openvino-model-downloader
            templateRef:
              name: openvino-model-downloader-wft
              template: openvino-model-downloader
---
{{- end -}}

templates/rollout.yaml

Argo Rollout workload: container spec, probes, resources and rollout strategy.

Show manifest
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: {{ .Release.Namespace }}-rollout
  namespace: {{ .Release.Namespace }}
  labels:
    app: {{ .Release.Namespace }}
spec:
  replicas: 1
  strategy:
    canary:
      steps:
      - setWeight: 25
      - pause: {duration: 30s}
      - setWeight: 50
      - pause: {duration: 45s}
      - setWeight: 100
  selector:
    matchLabels:
      app: {{ .Release.Namespace }}
  template:
    metadata:
      labels:
        app: {{ .Release.Namespace }}
    spec:
      containers:
      - name: {{ .Release.Namespace }}
        image: {{ .Values.image }}
        securityContext:
          privileged: true
          runAsUser: 0
          runAsGroup: 0
        env:
        - name: API_KEY
          valueFrom:
            secretKeyRef:
              name: openvino-secret
              key: OPENVINO_API_KEY
        - name: PERFORMANCE_HINT
          value: "LATENCY"
        - name: NUM_STREAMS
          value: "1"
        args:
          - "--config_path"
          - "/models/config.json"
          - "--rest_port"
          - "{{ .Values.port }}"
          - "--metrics_enable"
          - "--log_level"
          - "INFO"
        ports:
        - containerPort: {{ .Values.port }}
        {{- with .Values.resources }}
        resources:
          {{- toYaml . | nindent 10 }}
        {{ end }}
        volumeMounts:
        - name: models
          mountPath: /models
        - name: dri
          mountPath: /dev/dri
        - name: config
          mountPath: /models/config.json
          subPath: config.json
      dnsPolicy: ClusterFirst
      restartPolicy: Always
      schedulerName: default-scheduler
      volumes:
      - name: models
        hostPath:
          path: /mnt/ssd-large/genai-models/
          type: Directory
      - name: dri
        hostPath:
          path: /dev/dri
          type: Directory
      - name: config
        configMap:
          name: openvino-configmap

templates/service.yaml

ClusterIP Service fronting the workload.

Show manifest
apiVersion: v1
kind: Service
metadata:
  name: {{ .Release.Namespace }}-service
  namespace: {{ .Release.Namespace }}
spec:
  type: ClusterIP
  selector:
    app: {{ .Release.Namespace }}
  ports:
    - name: http
      protocol: TCP
      port: {{ .Values.port }}
      targetPort: {{ .Values.port }}

templates/workflow-template.yaml

Reusable Argo WorkflowTemplate for model-server tasks.

Show manifest
apiVersion: argoproj.io/v1alpha1
kind: WorkflowTemplate
metadata:
  name: openvino-model-downloader-wft
  namespace: {{ .Release.Namespace }}
spec:
  templates:
    - name: openvino-model-downloader
      serviceAccountName: argo-workflow-sa
      nodeSelector:
        kubernetes.io/hostname: {{ .Values.runOnHost}}
      script:
        image: {{ .Values.image }}
        securityContext:
          privileged: true
          runAsUser: 0
          runAsGroup: 0
        command: [bash]
        env:
        - name: HF_TOKEN
          valueFrom:
            secretKeyRef:
              name: openvino-secret
              key: HF_TOKEN
        source: |
          ARGS=(
            --pull
            --source_model "{{`{{workflow.parameters.modelRepo}}`}}"
            --model_repository_path /models
            --model_name "{{`{{workflow.parameters.modelName}}`}}"
            --task "{{`{{workflow.parameters.modelTask}}`}}"
            --target_device "{{`{{workflow.parameters.targetDevice}}`}}"
            --overwrite_models
          )


          # Text generation params
          if [ "{{`{{workflow.parameters.modelTask}}`}}" = "text_generation" ]; then
            if [ -n "{{`{{workflow.parameters.toolParser}}`}}" ]; then
              ARGS+=(--tool_parser "{{`{{workflow.parameters.toolParser}}`}}")
            fi
            if [ -n "{{`{{workflow.parameters.reasoningParser}}`}}" ]; then
              ARGS+=(--reasoning_parser "{{`{{workflow.parameters.reasoningParser}}`}}")
            fi
            if [ -n "{{`{{workflow.parameters.ggufFilename}}`}}" ]; then
              ARGS+=(--gguf_filename "{{`{{workflow.parameters.ggufFilename}}`}}")
            fi
          fi

          # Image generation params
          if [ "{{`{{workflow.parameters.modelTask}}`}}" = "image_generation" ]; then
            if [ -n "{{`{{workflow.parameters.defaultNumInferenceSteps}}`}}" ]; then
              ARGS+=(--default_num_inference_steps "{{`{{workflow.parameters.defaultNumInferenceSteps}}`}}")
            fi
          fi

          /ovms/bin/ovms "${ARGS[@]}"


          GRAPH=/models/{{`{{workflow.parameters.modelRepo}}`}}/graph.pbtxt
          echo "Looking for graph file in $GRAPH"

          if [ "{{`{{workflow.parameters.modelTask}}`}}" = "text_generation" ]; then
            if [ -f "$GRAPH" ]; then
              sed -i 's/max_num_seqs:256,/max_num_seqs:4,/' "$GRAPH"
              sed -i 's/cache_size: 0,/cache_size: 4,/' "$GRAPH"
              echo "Patched graph.pbtxt"
            else
              echo "WARNING: graph.pbtxt not found at $GRAPH"
            fi
          fi

        volumeMounts:
        - name: text-models-pv
          mountPath: /models
      {{- with .Values.tolerations }}
      tolerations:
        {{- toYaml . | nindent 8 }}
      {{ end }}
      volumes:
      - name: text-models-pv
        hostPath:
          path: /mnt/ssd-large/genai-models/
          type: Directory

Trade-offs

Decision. OpenVINO Model Server

Alternative. Ollama or vLLM

Why. One endpoint serves multiple model types, and the Intel runtime is optimised and fits the amd64 node.

← Back to AI Stack · All service groups