Free course · 8 parts · Intermediate

Local LLMs, hands-on: run models on your own machine

Run LLMs locally with Ollama, llama.cpp and LM Studio. Hardware and memory math, quantization, using local models in code, local RAG and hybrid setups.

Follow on YouTube

Course outline

  1. 1
    Why run locally, and your first model

    Privacy, cost and offline use, with a working model.

    Coming soon
  2. 2
    Hardware and memory math

    Will it fit? Do the arithmetic first.

    Coming soon
  3. 3
    Quantization and choosing a model

    Size, quality and speed trade-offs.

    Coming soon
  4. 4
    Ollama deep dive

    Modelfiles, the API and day-to-day use.

    Coming soon
  5. 5
    llama.cpp and LM Studio

    The engine under the tools.

    Coming soon
  6. 6
    Local LLMs in your code

    Use a local model from an app.

    Coming soon
  7. 7
    Local RAG

    Chat with your documents without the cloud.

    Coming soon
  8. 8
    Speed, evals and hybrid with Claude

    Measure it, then mix local and hosted.

    Coming soon

Run, measure and use models on your own hardware, in eight parts. Every part has a short concept video, then a hands-on video where we build it live. The written version of each part appears on this site as its video goes live, with code you can copy.

Frequently asked questions

Can I run an LLM on a laptop without a GPU?

Yes, small quantized models run on CPU. They are slower, so the course shows how to estimate speed and memory before you start.

Is a local LLM as good as a hosted one?

Usually not at the top end. The course covers when a local model is good enough and how to combine both.