Fundamentals

How to catch AI hallucinations in your code

You ask for code. It returns a clean function that calls response.json_strict(). That method does not exist. The model did not lie on purpose: it predicted the text that would most plausibly come next. This is discernment work, the third D of the 4D framework: you judge the output before you trust it.

The five shapes hallucinated code takes#

Wrong parameter or option namesusually fails loudlyeasy to catchInvented methods and flagscompiler or runtime errorOutdated APIsworked two versions agoPackages that do not exista supply-chain risk if someone registers the namePlausible but wrong logicruns fine, gives wrong answershardest tocatch
The scary ones are quiet. The code runs; it is just wrong.

Seven checks that catch it#

  1. Compile and type-check. Strict TypeScript, mypy, go vet. Invented methods fail instantly.
  2. Run the tests. If there are none, write one for the behavior you asked for. Tests are the only judge that does not guess.
  3. Look up every unfamiliar name. If you cannot find a function in the official docs, it probably does not exist.
  4. Verify packages before installing. Check the registry page, the publisher, the download count and the repo link. A package the AI named that you have never heard of deserves suspicion.
  5. Check the version. Ask which version the example targets, then compare it to what you run.
  6. Run it on real input. Edge cases, empty values, large values, Unicode.
  7. Read the diff. Ask "what would break if this assumption were wrong?" about each line you do not understand.

Prompt in ways that reduce it#

  • Provide the facts. Paste the real type definition, the schema, the error message, or the doc excerpt. The model stops guessing what it can read.
  • Give it permission to say no. "If you are not sure an API exists, say so instead of guessing."
  • Ask for sources it can quote. "Quote the relevant line from the docs I pasted."
  • Give it a verifier. An agent that can run the tests and read the failure corrects itself. See Claude Code tips.
  • Ask for a second opinion in a fresh chat. Paste the code and ask what is wrong with it. A fresh context catches things the first one defended.

The confident-tone trap#

The model does not say "hm, not sure". It says "Certainly! Here you go." Tone carries no information about correctness. Treat every answer as a draft from a fast colleague who never admits doubt.

A 60-second review routine#

bash
# 1. does it compile?
npx tsc --noEmit
# 2. do the tests pass?
npm test
# 3. does every new dependency exist and look legitimate?
git diff package.json

If all three are green and you have read the diff, you are in good shape. If you skipped any, you are vibe coding. See vibe coding vs AI-assisted engineering.

Frequently asked questions

What is an AI hallucination in code?

It is code that looks plausible but references things that do not exist or behave differently: an invented function, a wrong parameter, a made-up flag, or a package name that is not on the registry.

Why do LLMs hallucinate code?

A language model predicts likely text. When it lacks the exact fact, it still produces something that looks right. It is most likely to do this with niche libraries, new versions, and your private code it has never seen.

How do I check AI-generated code for hallucinations?

Compile or type-check it, run the tests, look up every unfamiliar function and package in the official docs, and run the code on real input before you trust it.

Do hallucinations go away with better models?

They get rarer, not extinct. Treat verification as part of the workflow, not as a fix for weak models.