Decision. OpenVINO Model Server
Alternative. Ollama or vLLM
Why. One endpoint serves multiple model types, and the Intel runtime is optimised and fits the amd64 node.
Self-hosted LLM, embedding, and image models behind an OpenAI-compatible API
OpenVINO Model Server hosts the local model stack: LLMs, embedding models, and image models behind a single OpenAI-compatible endpoint. Consumers switch models by name, not by deployment.
Excerpt from argocd-apps/values.yaml in the argocd-apps chart,
with annotations added for this site:
openvino:
application: true
project: applications
# false: model servers are memory-hungry; upgrades are scheduled
autoSync: false
The full openvino/values.yaml from the service's
own chart:
runOnHost: opi5-worker-5-genai
image: registry.opi5cluster.co.uk/openvino:2026.3-gpu
url: openvino.opi5cluster.co.uk
port: 8000
# Monitoring: OVMS serves Prometheus metrics on the REST port (`--metrics_enable`),
# scraped by the `monitoring` Alloy DaemonSet → Prometheus/Grafana.
volume:
storageClass: longhorn-ssd-large
size: 300Gi
accessModes: ReadWriteMany
resources:
requests:
cpu: 6
memory: 10Gi
modelsToDownload:
Qwen3.5-4B:
modelRepo: OpenVINO/Qwen3.5-4B-int4-ov
modelTask: text_generation
targetDevice: GPU
Qwen3.6-27B:
modelRepo: OpenVINO/Qwen3.6-27B-int4-ov
modelTask: text_generation
targetDevice: GPU
toolParser: qwen3coder
reasoningParser: qwen3
Qwen3-Embedding:
modelRepo: OpenVINO/Qwen3-Embedding-0.6B-int8-ov
modelTask: embeddings
targetDevice: GPU
Qwen3-Reranker:
modelRepo: OpenVINO/Qwen3-Reranker-0.6B-int8-ov
modelTask: rerank
targetDevice: GPU
FLUX:
modelRepo: OpenVINO/FLUX.1-schnell-int4-ov
modelTask: image_generation
targetDevice: CPU
defaultNumInferenceSteps: 4
tolerations:
- key: opi5.cluster/role
operator: Equal
value: data
effect: NoSchedule
env:
- name: IPEX_DEVICES
value: GPU
- name: DPCPP_DEVICE_TYPE
value: GPU templates/config.yaml OpenVINO Model Server config: served models and plugin settings.
apiVersion: v1
kind: ConfigMap
metadata:
name: openvino-configmap
namespace: {{ .Release.Namespace }}
data:
config.json: |
{
"model_config_list": [
{{- $models := .Values.modelsToDownload -}}
{{- $first := true -}}
{{- range $name, $value := $models }}
{{- if not $first }},{{ end }}
{
"config": {
"name": "openai/{{ $name }}",
"base_path": "/models/{{ $value.modelRepo }}"
}
}
{{- $first = false -}}
{{- end }}
]
} templates/external-secret.yaml ExternalSecret syncing the service credentials from Vault into the namespace.
# Create External Service to extract values from Doppler
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
name: {{ .Release.Namespace }}-external-secret
namespace: {{ .Release.Namespace }}
spec:
refreshInterval: 30s
secretStoreRef:
kind: ClusterSecretStore
name: vault-cluster-secret-store
target:
name: {{ .Release.Namespace }}-secret
creationPolicy: Owner
data:
- secretKey: OPENVINO_API_KEY
remoteRef:
key: genai
property: openvino-api-key
- secretKey: HF_TOKEN
remoteRef:
key: huggingface
property: token templates/http-route.yaml Gateway API HTTPRoute exposing the service through the Istio gateway under opi5cluster.co.uk.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: {{ .Release.Namespace }}-httproute
namespace: {{ .Release.Namespace }}
annotations:
link.argocd.argoproj.io/external-link: "https://{{ .Values.url }}"
spec:
parentRefs:
- name: istio-gateway
namespace: istio
sectionName: websecure
hostnames:
- {{ .Values.url }}
rules:
- backendRefs:
- name: {{ .Release.Namespace }}-service
port: {{ .Values.port }} templates/models-downloader-wf.yaml Argo Workflow downloading models into the OpenVINO model storage.
{{- $nameSpace := .Release.Namespace }}
{{- range $model, $property := .Values.modelsToDownload }}
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
name: {{ $model | lower }}-model-downloader-wf
spec:
entrypoint: openvino-model-downloader
serviceAccountName: argo-workflow-sa
arguments:
parameters:
- name: modelName
value: "{{ $model }}"
- name: modelRepo
value: "{{ $property.modelRepo | default "" }}"
- name: modelTask
value: "{{ $property.modelTask | default "" }}"
- name: targetDevice
value: "{{ $property.targetDevice | default "" }}"
- name: toolParser
value: "{{ $property.toolParser | default "" }}"
- name: reasoningParser
value: "{{ $property.reasoningParser | default "" }}"
- name: ggufFilename
value: "{{ $property.ggufFilename | default "" }}"
- name: defaultNumInferenceSteps
value: "{{ $property.defaultNumInferenceSteps | default "" }}"
templates:
- name: openvino-model-downloader
steps:
- - name: openvino-model-downloader
templateRef:
name: openvino-model-downloader-wft
template: openvino-model-downloader
---
{{- end -}} templates/rollout.yaml Argo Rollout workload: container spec, probes, resources and rollout strategy.
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: {{ .Release.Namespace }}-rollout
namespace: {{ .Release.Namespace }}
labels:
app: {{ .Release.Namespace }}
spec:
replicas: 1
strategy:
canary:
steps:
- setWeight: 25
- pause: {duration: 30s}
- setWeight: 50
- pause: {duration: 45s}
- setWeight: 100
selector:
matchLabels:
app: {{ .Release.Namespace }}
template:
metadata:
labels:
app: {{ .Release.Namespace }}
spec:
containers:
- name: {{ .Release.Namespace }}
image: {{ .Values.image }}
securityContext:
privileged: true
runAsUser: 0
runAsGroup: 0
env:
- name: API_KEY
valueFrom:
secretKeyRef:
name: openvino-secret
key: OPENVINO_API_KEY
- name: PERFORMANCE_HINT
value: "LATENCY"
- name: NUM_STREAMS
value: "1"
args:
- "--config_path"
- "/models/config.json"
- "--rest_port"
- "{{ .Values.port }}"
- "--metrics_enable"
- "--log_level"
- "INFO"
ports:
- containerPort: {{ .Values.port }}
{{- with .Values.resources }}
resources:
{{- toYaml . | nindent 10 }}
{{ end }}
volumeMounts:
- name: models
mountPath: /models
- name: dri
mountPath: /dev/dri
- name: config
mountPath: /models/config.json
subPath: config.json
dnsPolicy: ClusterFirst
restartPolicy: Always
schedulerName: default-scheduler
volumes:
- name: models
hostPath:
path: /mnt/ssd-large/genai-models/
type: Directory
- name: dri
hostPath:
path: /dev/dri
type: Directory
- name: config
configMap:
name: openvino-configmap templates/service.yaml ClusterIP Service fronting the workload.
apiVersion: v1
kind: Service
metadata:
name: {{ .Release.Namespace }}-service
namespace: {{ .Release.Namespace }}
spec:
type: ClusterIP
selector:
app: {{ .Release.Namespace }}
ports:
- name: http
protocol: TCP
port: {{ .Values.port }}
targetPort: {{ .Values.port }} templates/workflow-template.yaml Reusable Argo WorkflowTemplate for model-server tasks.
apiVersion: argoproj.io/v1alpha1
kind: WorkflowTemplate
metadata:
name: openvino-model-downloader-wft
namespace: {{ .Release.Namespace }}
spec:
templates:
- name: openvino-model-downloader
serviceAccountName: argo-workflow-sa
nodeSelector:
kubernetes.io/hostname: {{ .Values.runOnHost}}
script:
image: {{ .Values.image }}
securityContext:
privileged: true
runAsUser: 0
runAsGroup: 0
command: [bash]
env:
- name: HF_TOKEN
valueFrom:
secretKeyRef:
name: openvino-secret
key: HF_TOKEN
source: |
ARGS=(
--pull
--source_model "{{`{{workflow.parameters.modelRepo}}`}}"
--model_repository_path /models
--model_name "{{`{{workflow.parameters.modelName}}`}}"
--task "{{`{{workflow.parameters.modelTask}}`}}"
--target_device "{{`{{workflow.parameters.targetDevice}}`}}"
--overwrite_models
)
# Text generation params
if [ "{{`{{workflow.parameters.modelTask}}`}}" = "text_generation" ]; then
if [ -n "{{`{{workflow.parameters.toolParser}}`}}" ]; then
ARGS+=(--tool_parser "{{`{{workflow.parameters.toolParser}}`}}")
fi
if [ -n "{{`{{workflow.parameters.reasoningParser}}`}}" ]; then
ARGS+=(--reasoning_parser "{{`{{workflow.parameters.reasoningParser}}`}}")
fi
if [ -n "{{`{{workflow.parameters.ggufFilename}}`}}" ]; then
ARGS+=(--gguf_filename "{{`{{workflow.parameters.ggufFilename}}`}}")
fi
fi
# Image generation params
if [ "{{`{{workflow.parameters.modelTask}}`}}" = "image_generation" ]; then
if [ -n "{{`{{workflow.parameters.defaultNumInferenceSteps}}`}}" ]; then
ARGS+=(--default_num_inference_steps "{{`{{workflow.parameters.defaultNumInferenceSteps}}`}}")
fi
fi
/ovms/bin/ovms "${ARGS[@]}"
GRAPH=/models/{{`{{workflow.parameters.modelRepo}}`}}/graph.pbtxt
echo "Looking for graph file in $GRAPH"
if [ "{{`{{workflow.parameters.modelTask}}`}}" = "text_generation" ]; then
if [ -f "$GRAPH" ]; then
sed -i 's/max_num_seqs:256,/max_num_seqs:4,/' "$GRAPH"
sed -i 's/cache_size: 0,/cache_size: 4,/' "$GRAPH"
echo "Patched graph.pbtxt"
else
echo "WARNING: graph.pbtxt not found at $GRAPH"
fi
fi
volumeMounts:
- name: text-models-pv
mountPath: /models
{{- with .Values.tolerations }}
tolerations:
{{- toYaml . | nindent 8 }}
{{ end }}
volumes:
- name: text-models-pv
hostPath:
path: /mnt/ssd-large/genai-models/
type: Directory Decision. OpenVINO Model Server
Alternative. Ollama or vLLM
Why. One endpoint serves multiple model types, and the Intel runtime is optimised and fits the amd64 node.