https://llm.<domain>/v1. No tokens are metered by a vendor
and no prompt ever leaves the network. No consumer ever needs to know which
machine actually served the request.
The name resolves to a load-balanced pool of three stateless
LiteLLM routers. The routers hold the model table
and fan requests out by tier:
- Large tier: the biggest models run on an Apple-Silicon machine behind an authenticated TLS gate; the router holds the bearer token, consumers never see it. Which models are actually resident changes as the roster is re-evaluated. See What is actually loaded for the current, dated roster rather than a name hard-coded here. This is the primary tier: the router prefers it whenever it’s up.
- Fast tier: small, quick models (
qwen3-4band theembeddingsmodel) on a Radeon RX 6800 in a privileged LXC, served by llama-swap fronting llama.cpp’sllama-server(ROCm build, models co-resident). It now also backs up the large tier as a fallback leg. It’s fail-fast, not retrying: it accepts a request immediately if it has capacity or rejects immediately if it doesn’t, so a busy fast tier can never stall a caller waiting on a retry or backoff. - CPU standby: the same llama.cpp lineup, CPU-only, on a node that sleeps most of the day; cold capacity if both GPU tiers are out.
- A free hosted tier, reached through a stable handle whose underlying model lives in the provider’s own configuration rather than in anything committed here. See the stable-handle pattern. It sits behind every local tier in the fallback order: local capacity is always tried first.
- Paid providers, reached only by explicit model identifier, never as an implicit fallback. See Delegation: the cheapest capable tier.
How you reach it
Teal is a client, ink is the DNS and reverse-proxy edge, and coral is the local serving fabric. Dashed teal is a hosted tier reached only as fallback or by explicit request. Traefik health-checks the router pool (/health/liveliness) and drops a dead router without the endpoint changing.
Stopping one router live is a non-event. Every name resolves through
Technitium and carries the wildcard certificate, so it is HTTPS end to end.
What’s in the stack
All guests are DHCP-first LXCs on the
ai VLAN with deterministic MAC
reservations, so every address is a DNS name. llm-fast is a privileged LXC
with the GPU passed through (/dev/kfd + /dev/dri); models live on a 120 GB
fast-pool volume. The LXCs, firewall groups, and the Traefik pool entries are
provisioned by tofu-proxmox. llama.cpp,
llama-swap, the routers, and Open WebUI are configured by
ansible-proxmox-ai.
This is not the same “local AI” as the
Apple Silicon stack. That one is the MLX
server on this MacBook, tuned to hold one resident model for local-first
work. This page is the shared fabric: always on, LAN-wide, and the
MacBook falls back to it through the same router endpoint.
Use it from your Mac
Everything below is reachable by DNS name over HTTPS. Replaceexample.net
with your homelab’s internal domain. The router requires a bearer key (issued
per consumer), unlike the old unauthenticated single-backend endpoint.
1 · Browser
Openhttps://chat.example.net, sign in, pick a model, and chat. This is
the full Open WebUI: conversation history, system prompts, and file uploads,
all talking to the router like every other consumer.
2 · OpenAI-compatible API
Any OpenAI client works. Change the base URL and send the key:GET /v1/models lists every alias the fabric serves. Picking a model picks a
tier, and the router handles placement, fallbacks, and the heavy backend’s
authentication for you.
Related
tofu-proxmox
Provisions the LXCs, GPU passthrough, firewall groups, and the Traefik router pool.
ansible-proxmox-ai
Installs llama.cpp + llama-swap, the LiteLLM routers, and Open WebUI.
LXC vs Docker
Why the inference stack runs as native LXC, not Docker.
AI development pipeline
The other “local AI”: MLX on the workstation, local-first with fabric fallback.