Desperately Seeking LLMs: What Actually Works on an 8 GB Laptop GPU
A measurement-driven account of tuning Ollama and a small parade of local language models on a Ryzen laptop with an RTX 3070 8 GB—from unified-memory disasters and dense-model failures to MoE, MTP, context cliffs, and the models that finally worked.
Read article : Desperately Seeking LLMs: What Actually Works on an 8 GB Laptop GPU


