<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Squirrel Glitch</title><description>Working notes on local LLM infrastructure, agent pipelines, and shop hardware — measured, with the data.</description><link>https://squirrelglitch.com/</link><language>en-us</language><item><title>Three vLLM models on one RTX 3090, measured</title><link>https://squirrelglitch.com/research/3090-bake-off/</link><guid isPermaLink="true">https://squirrelglitch.com/research/3090-bake-off/</guid><description>Qwen3.8-27B, Ornith-1.5-35B-A3B and DiffusionGemma-26B-A4B, each booted through llama-swap on the same workstation, timed with the same harness, and judged by whether the code they returned actually runs. With a small-model round and a controlled prompt-and-sampling re-test.</description><pubDate>Sun, 30 Aug 2026 12:00:00 GMT</pubDate><category>local-llm</category><category>vllm</category><category>llama-swap</category><category>benchmark</category></item><item><title>Putting a coding model on the RTX 3090</title><link>https://squirrelglitch.com/research/3090-local-coding-stack/</link><guid isPermaLink="true">https://squirrelglitch.com/research/3090-local-coding-stack/</guid><description>What is actually worth running on a dedicated 24 GB Ampere card in late August 2026, measured rather than advertised: llama.cpp + llama-swap against vLLM W4A16, the Ampere-specific traps, and how it wires into the tools already running on the workstation.</description><pubDate>Sun, 30 Aug 2026 12:00:00 GMT</pubDate><category>local-llm</category><category>llama.cpp</category><category>vllm</category><category>ampere</category></item></channel></rss>