Real tests with recorded environments, repeatable commands, measured results and explicit limitations.
What happens when a 4B Q4 model is larger at runtime than the available GPU budget? We measured partial CPU/GPU offload at three context sizes and tested whether freeing Windows VRAM changed placement.
Can a 3 GB Pascal GPU run a current local LLM entirely on the GPU? We installed Ollama on Windows, recorded placement and VRAM snapshots, then ran three API measurements.