Self-hosted LLM using Ollama with NVIDIA GPU support and Open WebUI.
  • Shell 58.7%
  • Dockerfile 17.7%
  • Python 15.7%
  • Jinja 7.9%
Find a file
2026-09-21 06:22:06 -05:00
.github feat(m1): baseline & quick wins — 36→56 tps via 32k context 2026-05-05 16:20:59 -05:00
autocomplete chore: commit autocomplete CUDA phi2, compose tweaks, and imagegen plan draft 2026-09-20 22:12:48 -05:00
coding chore: commit autocomplete CUDA phi2, compose tweaks, and imagegen plan draft 2026-09-20 22:12:48 -05:00
docs chore: commit autocomplete CUDA phi2, compose tweaks, and imagegen plan draft 2026-09-20 22:12:48 -05:00
imagegen fix: use named volume instead of bind mount for imagegen models\n\nModels copied into Docker volume at docker-data/volumes/.../_data.\nRemoves bind mount to /home/james/Games/comfyui-models. 2026-09-21 06:22:06 -05:00
infra feat: add imagegen category — ComfyUI + Qwen-Image-2.1 GGUF 2026-09-21 05:57:01 -05:00
logs chore: update monitoring submodule (tps-monitor simplify) 2026-06-30 14:22:44 -05:00
multimodal/audio-sfx/__pycache__ perf(m3-s1): llama.cpp updated to 94a220c — median 84-86 -> 89.9 tps (+4-6%) 2026-06-03 15:24:14 -05:00
scripts feat: add imagegen category — ComfyUI + Qwen-Image-2.1 GGUF 2026-09-21 05:57:01 -05:00
thunder-llm-monitoring@1cc6668012 chore: commit autocomplete CUDA phi2, compose tweaks, and imagegen plan draft 2026-09-20 22:12:48 -05:00
.env docs: add phase 3 milestone 2 vLLM MXFP4 deployment docs and scripts 2026-06-30 08:20:36 -05:00
.gitignore chore: consolidate ROCm as default, add compose profiles, update docs 2026-07-26 20:57:45 -05:00
.gitmodules refactor: split into multi-category layout (S1 — coding category) 2026-08-07 07:19:47 -05:00
CLAUDE.md infra: move submodules into docker/ and fix GPU device stability 2026-08-06 08:49:32 -05:00
llamacpp-cuda-phi2.Dockerfile chore: commit autocomplete CUDA phi2, compose tweaks, and imagegen plan draft 2026-09-20 22:12:48 -05:00
README.md feat: add imagegen category — ComfyUI + Qwen-Image-2.1 GGUF 2026-09-21 05:57:01 -05:00

Thunder LLM

A self-hosted LLM server instance for use with Pi and other AI coding agents.

Purpose: Reliable, high-throughput local inference for development workflows. Server endpoint: http://localhost:8080 (llama.cpp, ROCm backend)


Server

Current configuration:

Component Value
Model Qwen3.6-27B-Q6_K (GGUF)
Hardware AMD RX 9700 (32 GB GDDR6)
Backend llama.cpp (Vulkan, --reasoning on, MTP spec decoding)
Build b9209
Context 256K
Speculative decoding draft-MTP (MTP speculative decoding)
Reasoning --reasoning on, budget 16384, client-controlled via chat_template_kwargs
Server default --chat-template-kwargs '{"preserve_thinking": true}' (server-side default for thinking preservation)
Throughput ~99 tps decode (Vulkan)
Container Docker, managed via docker compose

Health: curl -s http://localhost:8080/health


Available Services

Coding (port 8080)

Each compose file in coding/compose/ defines a single service with a permutation name: <backend>-<device>-<model-spec>. Switch between them with ai-server coding <engine>.

Service Backend Model Notes
llamacpp-rocm-qwen36-35b-q5 llama.cpp ROCm Qwen3.6-35B-A3B-UD-Q5_K_XL Fallback — ~66 tps, no GUI impact
llamacpp-rocm-qwen36-27b-q6 llama.cpp ROCm Qwen3.6-27B-Q6_K 27B dense model on ROCm
llamacpp-rocm-qwen36-27b-q6-mtp llama.cpp ROCm (MTP) Qwen3.6-27B-Q6_K 27B with draft-MTP speculative decoding
llamacpp-rocm-qwen36-27b-q4 llama.cpp ROCm Qwen3.6-27B-Q4 Smallest 27B variant
llamacpp-vulkan-qwen36-35b-q5 llama.cpp Vulkan Qwen3.6-35B-A3B-UD-Q5_K_XL Faster decode (~99 tps) but GUI sluggish during heavy PP
llamacpp-vulkan-qwen36-27b-q6 llama.cpp Vulkan Qwen3.6-27B-Q6_K 27B dense on Vulkan
llamacpp-vulkan-qwen36-27b-q6-mtp llama.cpp Vulkan (MTP) Qwen3.6-27B-Q6_K Default daily driver — ~99 tps, MTP speculative decoding
vllm-rocm-qwen36-35b-q5 vLLM ROCm Qwen3.6-35B-A3B-UD-Q5_K_XL Standard vLLM ROCm
vllm-rocm-qwen36-35b-mxfp4 vLLM ROCm (MXFP4) Qwen3.6-35B-A3B-MXFP4 RDNA4-specific kernels, custom Dockerfile
vllm-rocm-qwen36-35b-kyuz0 vLLM ROCm (patched) Qwen3.6-35B-A3B-MXFP4 kyuz0's gfx1201 patch with AITER unified attention
vllm-rocm-qwen36-35b-tcclaviger vLLM ROCm (patched) Qwen3.6-35B-A3B-MXFP4 tcclaviger's 1,980 AMD-specific patches

Why Vulkan MTP is the default: Vulkan provides faster decode throughput (~99 tps vs ~66 tps ROCm). The MTP speculative decoding further boosts effective speed. ROCm remains available as a fallback when GPU display contention is a concern.

Autocomplete (port 8081)

Ghost-text autocomplete for code editors. Managed independently from the coding server — can run alongside it.

Service Backend Model Notes
llamacpp-rocm-qwen25-coder-15b-q4 llama.cpp ROCm Qwen2.5-Coder-1.5B-Q4_K_M Small, fast completions (~980MB)

CLI: ai-server ac <engine> to switch, ai-server ac status to check health.

Image Generation (port 8188)

ComfyUI with GGUF diffusion models. Cannot run simultaneously with the coding server — both compete for the 32 GB VRAM. The switch script stops the coding server automatically.

Service Backend Model Notes
comfyui-rocm ComfyUI + ComfyUI-GGUF (ROCm PyTorch) Qwen-Image-2.1 Q4_K_M GGUF Primary — uses ROCm PyTorch for diffusion
comfyui-vulkan ComfyUI + ComfyUI-GGUF (CPU PyTorch + Vulkan) Qwen-Image-2.1 Q4_K_M GGUF Test — CPU PyTorch, Vulkan via llama.cpp

Models (stored on Games drive, ~10.5 GB total):

Model File Size Purpose
Diffusion (DiT) qwen_image_2.1_Q4_K_M.gguf 4.0 GB Core image generation
Text Encoder qwen3vl_8b_w4a8.safetensors 5.9 GB Required — DiT embedding space is fixed
VAE qwen_image_2.1_vae_bf16.safetensors 645 MB Latent space decode

CLI: ai-server img <engine> to switch (stops coding server), ai-server img status, ai-server img down.

⚠️ VRAM constraint: Coding server (~14 GB VRAM) + image gen (~15+ GB VRAM) > 32 GB available. Must stop one before starting the other.


Container Health Monitoring

An idle-aware monitor automatically checks TPS and restarts the coding container when performance degrades. Prevents the ~30–40 tps slowdown that can occur after 6–8 hours of continuous uptime.

How it works:

  • Runs every 2 hours via cron (0 */2 * * *)
  • Benchmarks the container (5 warm-up runs + 3 measured, median TPS)
  • Logs TPS on every check — active, idle, fresh, old
  • Only restarts when: idle (no requests in 10 min) + uptime ≥ 4 hours + TPS < 70
  • Never restarts while you're actively using the service
  • Logs: logs/monitor.log
  • Monitor binary: thunder-llm-monitoring submodule (AOT-compiled .NET 10.0)

Manual check:

/home/james/.thunder-llm/Thunder.Llm.Monitor tps-monitor

Monitoring

Two subsystems track server health and GPU telemetry:

GPU Wattage Monitor (systemd)

Continuous GPU telemetry — 1-second resolution logging to ~/.thunder-llm/logs/wattage-YYYY-MM-DD.csv:

  • timestamp, watts, temp_edge_c, temp_junction_c, fan_pct, gpu_use_pct, vram_pct
  • Service: monitor-wattage.service (enabled, runs at boot)
  • Log: sudo journalctl -u monitor-wattage -f

TPS Health Monitor (cron)

Every 2 hours: benchmarks the LLM server (5 warmup + 3 measured, median TPS). If median TPS < 70 and container is ≥ 4 hours old, restarts the container.

  • Log: ~/.thunder-llm/logs/tps-YYYY-MM-DD.csv
  • Old files are downsampled daily (10-second intervals) to save space

Manual run:

/home/james/.thunder-llm/Thunder.Llm.Monitor tps-monitor

Monitoring Source

The monitor binary is an AOT-compiled .NET 10.0 app in the thunder-llm-monitoring submodule:

cd thunder-llm-monitoring && ./deploy.sh

Bench CLI

The bench CLI has been moved to a dedicated repository: thunder-llm-bench. See code.thundersizzle.tech/Thunder/thunder-llm-bench for the full CLI, docs, and history.

Quick benchmark (server-side):

# Raw TPS benchmark — no dependencies, just curl against the server
bash docs/work/bench-llamacpp.sh [warmups] [measured]

Quick Start

Global CLI

The ai-server command is available system-wide (symlinked to scripts/ai-server). It dispatches to category-specific scripts:

ai-server status                    # show status of all categories
ai-server coding <engine>           # switch coding engine (with rollback)
ai-server coding --force            # force restart coding engine
ai-server coding status             # coding health check
ai-server coding logs               # follow coding logs
ai-server coding engines            # list available coding engines
ai-server coding down               # stop all coding engines
ai-server ac <engine>               # switch autocomplete engine (with rollback)
ai-server ac --force                # force restart autocomplete engine
ai-server ac status                 # autocomplete health check
ai-server ac logs                   # follow autocomplete logs
ai-server ac engines                # list available autocomplete engines
ai-server ac down                   # stop all autocomplete engines
ai-server img <engine>              # switch image generation engine (stops coding server)
ai-server img --force               # force restart image generation engine
ai-server img status                # image generation health check
ai-server img logs                  # follow image generation logs
ai-server img engines               # list available image generation engines
ai-server img down                  # stop all image generation engines

Run the default coding server (ROCm)

ai-server coding llamacpp-rocm-qwen36-35b-q5

Run an alternative coding backend

ai-server coding llamacpp-vulkan-qwen36-35b-q5    # Vulkan backend
ai-server coding vllm-rocm-qwen36-35b-kyuz0       # vLLM RDNA4 experimental

Start autocomplete

ai-server ac llamacpp-rocm-qwen25-coder-15b-q4

Note: Each compose file defines exactly one service. The ai-server command wraps category-specific switch-engine.sh scripts with automatic health checks and rollback on failure.

Run a quick benchmark (server-side)

# Raw TPS benchmark — no build step, just curl against the running server
bash docs/work/bench-llamacpp.sh 5 3

Full bench CLI

For the complete benchmarking suite (TPS, context probes, agentic evaluation, model discovery, etc.), use the dedicated thunder-llm-bench repo:

git clone ssh://git@code.thundersizzle.tech:2222/Thunder/thunder-llm-bench.git
cd thunder-llm-bench
dotnet build -c Release src/Thunder.LlmBench/Thunder.LlmBench.csproj
alias bench="dotnet run --project src/Thunder.LlmBench/ --"

Repository Structure

.
├── infra/                           # Shared build assets
│   ├── DockerFiles/                 # Dockerfiles (shared across categories)
│   │   ├── llamacpp-stable-roc.Dockerfile
│   │   ├── llamacpp-stable-vulkan.Dockerfile
│   │   ├── llamacpp-rocm.Dockerfile
│   │   ├── llamacpp-vulkan.Dockerfile
│   │   ├── llamacpp-unstable-roc.Dockerfile
│   │   ├── llamacpp-unstable-vulkan.Dockerfile
│   │   ├── vllm-rocm.Dockerfile
│   │   ├── vllm-mxfp4.Dockerfile
│   │   ├── comfyui-rocm.Dockerfile
│   │   └── comfyui-vulkan.Dockerfile
│   └── third-party/                 # LLaMA.cpp submodule (build dependency)
│       ├── llama.cpp-stable/
│       └── llama.cpp-unstable/
│
├── coding/                          # Main chat inference (port 8080)
│   ├── docker-compose.yaml          # name: thunder-llm
│   ├── compose/                     # 11 engine configs
│   │   ├── llamacpp-rocm-*.yaml
│   │   ├── llamacpp-vulkan-*.yaml
│   │   └── vllm-rocm-*.yaml
│   ├── scripts/
│   │   ├── compose-wrapper.sh       # Resolves GPU device paths
│   │   └── switch-engine.sh         # Engine switching with rollback
│   └── qwen3.6/                     # Chat template
│
├── autocomplete/                    # Ghost-text autocomplete (port 8081)
│   ├── docker-compose.yaml          # name: thunder-llm-ac
│   ├── compose/
│   │   └── llamacpp-rocm-qwen25-coder-15b-q4.yaml
│   └── scripts/
│       ├── compose-wrapper.sh
│       └── switch-engine.sh
│
├── imagegen/                        # ComfyUI image generation (port 8188)
│   ├── docker-compose.yaml          # name: thunder-imagegen
│   ├── compose/
│   │   ├── comfyui-rocm.yaml        # ROCm PyTorch backend
│   │   └── comfyui-vulkan.yaml      # CPU PyTorch + Vulkan
│   └── scripts/
│       ├── compose-wrapper.sh
│       └── switch-engine.sh
│
├── scripts/ai-server                # Global CLI (symlink to /usr/local/bin/ai-server)
├── scripts/load-test.py             # Load testing / stress testing
├── thunder-llm-monitoring/          # GPU wattage + TPS monitoring (submodule)
│   ├── deploy.sh                    # AOT publish, systemd deploy, cron setup
│   └── src/Thunder.Llm.Monitor/     # Monitor source (AOT .NET 10.0)
├── docs/work/                       # Milestone docs, benchmarks, status
│   ├── STATUS.md                    # Current phase/milestone state
│   ├── infrastructure-split-plan.md # Multi-category split plan
│   ├── v1/phase1-.../               # Phase and milestone documentation
│   └── bench-llamacpp.sh            # Quick TPS benchmark script
├── logs/monitor.log                 # Health monitor logs
└── README.md                        # This file

Note: The bench CLI (Thunder.LlmBench) lives in a separate repo: thunder-llm-bench. This repo contains only the server infrastructure (Dockerfiles, compose, scripts, and server docs).


Container Management

Each compose file defines exactly one service. Use the ai-server CLI for all operations:

ai-server coding llamacpp-rocm-qwen36-35b-q5    # switch coding to ROCm
ai-server coding logs                           # view coding logs
ai-server ac llamacpp-rocm-qwen25-coder-15b-q4  # start autocomplete
ai-server img comfyui-rocm                       # start image generation (stops coding)

Category isolation

Each category (coding, autocomplete) is an independent Docker Compose project:

  • coding/ → project thunder-llm (port 8080)
  • autocomplete/ → project thunder-llm-ac (port 8081)
  • imagegen/ → project thunder-imagegen (port 8188)

Stopping one category does not affect the other. They share the GPU but have independent lifecycles.

Switching between backends (REQUIRED)

Always use ai-server for switching between backends. It wraps category-specific switch-engine.sh scripts, which auto-detect the running engine, health-check the new one, and roll back on failure:

  • Coding rolls back to llamacpp-rocm-qwen36-35b-q5
  • Autocomplete rolls back to llamacpp-rocm-qwen25-coder-15b-q4

Never call docker compose down/up to switch backends. Use it only for rebuilding containers, config changes, or volume operations — or if the switch script itself fails.

GPU Device Mapping

The AMD GPU's render node (/dev/dri/renderD*) can change number between reboots (e.g., renderD128 ↔ renderD129) depending on kernel device discovery order. Each category has its own compose-wrapper.sh that resolves the device at launch time:

/dev/dri/by-path/pci-0000:03:00.0-render → renderD128 (or whatever it is today)

The wrapper script resolves this symlink before Docker Compose parses the files, so the correct device is always passed regardless of kernel-assigned numbering.

After a reboot: if the renderD numbers swapped, the container's auto-restart (restart: unless-stopped) may attach the wrong device. Fix it with:

ai-server coding --force <current-engine>    # or: ai-server ac --force <engine>

This stops and recreates the container with the freshly-resolved device path.

Compose wrapper

Each category has its own compose-wrapper.sh (coding/scripts/, autocomplete/scripts/). It:

  1. Resolves AMD_RENDER_DEVICE from the PCI-stable symlink
  2. Changes to the category directory
  3. Passes all arguments through to docker compose

Both switch-engine.sh and the global ai-server CLI use it internally.

GPU memory

rocm-smi --showmeminfo vram

Reasoning / Thinking Mode

The server runs with --reasoning on and --reasoning-budget 16384 by default. The server also sets --chat-template-kwargs '{"preserve_thinking": true}' as a server-side default, so thinking output is preserved even when clients request non-reasoning mode (clients can still toggle thinking on/off).

How clients control reasoning:

  • Pi harness — uses thinkingFormat: "qwen-chat-template" in ~/.pi/agent/models.json. Specifies llamacpp/Qwen3.6-35B-A3B-Q5_K_M:high (or :off, :low, :medium, :xhigh) in the model pattern. Automatically injects {enable_thinking: true/false, preserve_thinking: true}.

  • Open WebUI (llm.thundersizzle.tech) — always uses reasoning mode (no client-side toggle). The pi harness hides thinking output via hideThinkingBlock: true.

  • Other clients — inject {"chat_template_kwargs": {"enable_thinking": true, "preserve_thinking": true}} for reasoning, or {"chat_template_kwargs": {"enable_thinking": false}} for non-reasoning mode. The server's preserve_thinking: true default still applies.

Adjusting reasoning budget:

The server default is 16384 tokens. For faster responses on simple queries, reduce in docker-compose.yaml:

--reasoning on
--reasoning-budget 8192  # lower = faster, still good for agentic work

Troubleshooting

TPS degraded to ~30–40

Container state degradation after extended uptime. The idle-aware monitor handles this automatically (every 2 hours, if idle + ≥ 4h uptime + TPS < 70). Manual restart:

ai-server coding --force llamacpp-rocm-qwen36-35b-q5

Shader cache stale after Mesa upgrade

Clear and rebuild:

rm -rf ~/.cache/mesa_shader_cache/
ai-server coding --force llamacpp-vulkan-qwen36-35b-q5

(Only applies to Vulkan backend — ROCm does not use Mesa shader caching.)

Model returning garbage tokens ("0", "&!", "3 years ago")

Caused by --reasoning off combined with --chat-template-kwargs conflict, or corrupted conversation context. Fix:

  1. Ensure server runs --reasoning on (not --reasoning off)
  2. The server defaults to --chat-template-kwargs '{"preserve_thinking": true}' — this is intentional. Clients control reasoning via enable_thinking.
  3. Start a fresh conversation in the client (corrupted context from previous bad requests propagates)

Container won't start after reboot (GPU device mismatch)

The AMD GPU's render node number can change between reboots. If the container fails to start or returns errors about GPU access:

ai-server coding --force <current-engine>    # e.g., ai-server coding --force llamacpp-rocm-qwen36-35b-q5

This resolves the current PCI symlink and recreates the container with the correct device.

vLLM backends failing to start

The vLLM RDNA4 backends are experimental and may fail depending on ROCm driver version. Check logs:

docker logs vllm-rocm-qwen36-35b-kyuz0 -f
docker logs vllm-rocm-qwen36-35b-tcclaviger -f

Previous Hardware

NVIDIA RTX 3090 (24 GB) — v0.x phase. Migration to AMD RX 9700 is complete.

ROCm is the recommended daily driver for the R9700. Vulkan remains available for benchmarking and when maximum decode speed is needed.