Updated on Sep 6, 2026
• Vision: Qwen3‑VL now default; GPU→CPU fallback; improved model search + one‑tap mmproj.
• n_ctx auto‑expands; over‑limit inputs return a clean server error.
• Engine: llama.cpp b10621 / ggml 0.22 with broader model/GPU support.
• Defaults: Chat = Qwen3.5‑2B, Vision = Qwen3‑VL.
• Fixes: missing BOS; streaming continues.
