AI Stack
Local inference. No cloud API keys in the request path.
Every service in this group has its own page documenting how it runs: the values it is deployed with, and the trade-offs behind those choices.
| Service | What it does |
|---|---|
| OpenVINO Model Server | Self-hosted LLM, embedding, and image models behind an OpenAI-compatible API |
| Open WebUI | ChatGPT-style interface over the local models, with per-user history and presets |
| Qdrant | Vector database powering semantic search and grounding the local models with RAG |