build1 publisher
A 4 GB laptop GPU decodes quantised Gemma 4 at 4.27x the CPU rate on 1598 MiB
Two arms on the same laptop differ by one flag. The small card wins because llama.cpp leaves Gemma 4's 1.93 GB per-layer embedding table in mmap and pulls a few rows per token, so only about 1.08 GB of body is resident.
Publishers:dev.to
Reality
- Evidence64
- Adoption12
- Hype gap+5
- Incentives20
- Confidence58