E—DOCEO
AILearnTutorialsToolsLab
About
E—DOCEO
AILearnTutorialsToolsLabAbout
  1. E—DOCEO
  2. DOCEO LAB
E—DOCEO · TEST

DOCEO LAB

Real tests with recorded environments, repeatable commands, measured results and explicit limitations.

VERIFIEDTEST #002 · SEP 18, 2026

Qwen3 4B on a GTX 1060 3GB

What happens when a 4B Q4 model is larger at runtime than the available GPU budget? We measured partial CPU/GPU offload at three context sizes and tested whether freeing Windows VRAM changed placement.

12.16tok/s · context 4096
43–47%GPU placement
3.1–3.5GB runtime size
Open test record →
VERIFIEDTEST #001 · SEP 18, 2026

Gemma 3 1B on a GTX 1060 3GB

Can a 3 GB Pascal GPU run a current local LLM entirely on the GPU? We installed Ollama on Windows, recorded placement and VRAM snapshots, then ran three API measurements.

66.99median tok/s
100%GPU placement
3 GBGTX 1060 VRAM
Open test record →
E—DOCEO
AI · SOFTWARE · KNOWLEDGE

Knowledge you can use.

AILocal AIVRAMOllama
ExploreLearnTutorialsLab
E—DOCEOAboutMethodologyEditorial Policy
© 2026 E—DOCEO. Independent guides, practical tests and tools.