Testing Qwen3.8-27B KV cache dtypes
Qwen3.8-27B FP16: Is there a quality difference between FP8 and FP16 for the KV cache?
Local LLMvLLMDeveloper toolsQwen3.8
Field notes from computers, local AI, scientific software and the physics problems that send me back to the terminal.
Latest writing
Qwen3.8-27B FP16: Is there a quality difference between FP8 and FP16 for the KV cache?
The model, memory and gateway configuration behind the long-context language-model service running on mlrig.
A technical overview of the GPU compute host used for local language models, numerical experiments and isolated rental workloads.
Notes from configuring Roo Code with locally hosted DeepSeek and GPT-OSS models through vLLM.
Active projects