Testing Qwen3.8-27B KV cache dtypes
Qwen3.8-27B FP16: Is there a quality difference between FP8 and FP16 for the KV cache?
Chronological record
Experiments, benchmarks, failures and findings. Newest first; corrections stay visible.
15 entries
Qwen3.8-27B FP16: Is there a quality difference between FP8 and FP16 for the KV cache?
The model, memory and gateway configuration behind the long-context language-model service running on mlrig.
A technical overview of the GPU compute host used for local language models, numerical experiments and isolated rental workloads.
Notes from configuring Roo Code with locally hosted DeepSeek and GPT-OSS models through vLLM.
A working note on the Porter-Thomas distribution and fluctuations in reduced transition strengths.
Comparing orbital and spin contributions to calculated M1 strength functions.
Removing selected one-body transition-density contributions to identify which orbitals produce the low-energy enhancement.
Investigating why apparent excitation contributions exceed decay contributions in OBTD grids.
Tracing how KSHELL transition states are ordered when OBTD data is loaded.
Visualising orbital contributions in one-body transition densities for selected M1 transitions.
Notes about pygmy resonances and the scissors mode in nuclear physics.
Working notes about one-body transition operators, one-body transition densities, and second quantisation.
An overview of effective interactions and their role in nuclear shell-model calculations.
Mapping a symmetric shell-model Hamiltonian onto GPU threads reduced a 186-second CPU calculation to 1.26 seconds.
No entries match this filter.