Skip to main content
The homelab runs its own ChatGPT, and every other model size below it, behind one endpoint: https://llm.<domain>/v1. No tokens are metered by a vendor and no prompt ever leaves the network. No consumer ever needs to know which machine actually served the request. The name resolves to a load-balanced pool of three stateless LiteLLM routers. The routers hold the model table and fan requests out by tier:
  • Large tier: the biggest models run on an Apple-Silicon machine behind an authenticated TLS gate; the router holds the bearer token, consumers never see it. Which models are actually resident changes as the roster is re-evaluated. See What is actually loaded for the current, dated roster rather than a name hard-coded here. This is the primary tier: the router prefers it whenever it’s up.
  • Fast tier: small, quick models (qwen3-4b and the embeddings model) on a Radeon RX 6800 in a privileged LXC, served by llama-swap fronting llama.cpp’s llama-server (ROCm build, models co-resident). It now also backs up the large tier as a fallback leg. It’s fail-fast, not retrying: it accepts a request immediately if it has capacity or rejects immediately if it doesn’t, so a busy fast tier can never stall a caller waiting on a retry or backoff.
  • CPU standby: the same llama.cpp lineup, CPU-only, on a node that sleeps most of the day; cold capacity if both GPU tiers are out.
  • A free hosted tier, reached through a stable handle whose underlying model lives in the provider’s own configuration rather than in anything committed here. See the stable-handle pattern. It sits behind every local tier in the fallback order: local capacity is always tried first.
  • Paid providers, reached only by explicit model identifier, never as an implicit fallback. See Delegation: the cheapest capable tier.
The router chooses among the tiers it’s allowed to reach for a given request. It uses a learned quality-vs-cost tradeoff rather than a fixed order. See Cost as data. Each tier still registers an honest cost. In practice the order preceding is what falls out: local primary, then local GPU, then the free hosted tier, then paid providers.

How you reach it

Teal is a client, ink is the DNS and reverse-proxy edge, and coral is the local serving fabric. Dashed teal is a hosted tier reached only as fallback or by explicit request. Traefik health-checks the router pool (/health/liveliness) and drops a dead router without the endpoint changing. Stopping one router live is a non-event. Every name resolves through Technitium and carries the wildcard certificate, so it is HTTPS end to end.

What’s in the stack

All guests are DHCP-first LXCs on the ai VLAN with deterministic MAC reservations, so every address is a DNS name. llm-fast is a privileged LXC with the GPU passed through (/dev/kfd + /dev/dri); models live on a 120 GB fast-pool volume. The LXCs, firewall groups, and the Traefik pool entries are provisioned by tofu-proxmox. llama.cpp, llama-swap, the routers, and Open WebUI are configured by ansible-proxmox-ai.
This is not the same “local AI” as the Apple Silicon stack. That one is the MLX server on this MacBook, tuned to hold one resident model for local-first work. This page is the shared fabric: always on, LAN-wide, and the MacBook falls back to it through the same router endpoint.

Use it from your Mac

Everything below is reachable by DNS name over HTTPS. Replace example.net with your homelab’s internal domain. The router requires a bearer key (issued per consumer), unlike the old unauthenticated single-backend endpoint.

1 · Browser

Open https://chat.example.net, sign in, pick a model, and chat. This is the full Open WebUI: conversation history, system prompts, and file uploads, all talking to the router like every other consumer.

2 · OpenAI-compatible API

Any OpenAI client works. Change the base URL and send the key:
GET /v1/models lists every alias the fabric serves. Picking a model picks a tier, and the router handles placement, fallbacks, and the heavy backend’s authentication for you.

tofu-proxmox

Provisions the LXCs, GPU passthrough, firewall groups, and the Traefik router pool.

ansible-proxmox-ai

Installs llama.cpp + llama-swap, the LiteLLM routers, and Open WebUI.

LXC vs Docker

Why the inference stack runs as native LXC, not Docker.

AI development pipeline

The other “local AI”: MLX on the workstation, local-first with fabric fallback.