Free course · 8 parts · Intermediate
Local LLMs, hands-on: run models on your own machine
Run LLMs locally with Ollama, llama.cpp and LM Studio. Hardware and memory math, quantization, using local models in code, local RAG and hybrid setups.
Course outline
- 1Why run locally, and your first modelComing soon
Privacy, cost and offline use, with a working model.
- 2Hardware and memory mathComing soon
Will it fit? Do the arithmetic first.
- 3Quantization and choosing a modelComing soon
Size, quality and speed trade-offs.
- 4Ollama deep diveComing soon
Modelfiles, the API and day-to-day use.
- 5llama.cpp and LM StudioComing soon
The engine under the tools.
- 6Local LLMs in your codeComing soon
Use a local model from an app.
- 7Local RAGComing soon
Chat with your documents without the cloud.
- 8Speed, evals and hybrid with ClaudeComing soon
Measure it, then mix local and hosted.
Run, measure and use models on your own hardware, in eight parts. Every part has a short concept video, then a hands-on video where we build it live. The written version of each part appears on this site as its video goes live, with code you can copy.
Frequently asked questions
Can I run an LLM on a laptop without a GPU?
Yes, small quantized models run on CPU. They are slower, so the course shows how to estimate speed and memory before you start.
Is a local LLM as good as a hosted one?
Usually not at the top end. The course covers when a local model is good enough and how to combine both.