How to get reliable JSON from an LLM (structured output)
Your code needs data, not prose. The model gives you a friendly paragraph with JSON somewhere in the middle. Here are four ways to deal with that, ranked from weakest to strongest.
1 and 2: prompting for JSON#
Extract the fields below and reply with ONLY a JSON object, no other text:
{"name": string, "email": string, "plan": "free" | "pro" | "enterprise"}Fine for a quick script. In production you will meet the failures: a code fence around the JSON, a "Sure! Here it is:" before it, a trailing comma, a missing field. Each one is a crash in JSON.parse.
3 and 4: let the API enforce the shape#
Define the schema once in Zod, and let the SDK convert it and parse the reply into a typed object.
import Anthropic from "@anthropic-ai/sdk";
import { z } from "zod";
import { zodOutputFormat } from "@anthropic-ai/sdk/helpers/zod";
const Contact = z.object({
name: z.string(),
email: z.string(),
plan: z.enum(["free", "pro", "enterprise"]),
interests: z.array(z.string()),
demo_requested: z.boolean(),
});
const client = new Anthropic();
const response = await client.messages.parse({
model: "claude-opus-5-5",
max_tokens: 2048,
messages: [
{
role: "user",
content:
"Extract: Jane Doe (jane@co.com) wants Enterprise, interested in the API and SDKs, wants a demo.",
},
],
output_config: { format: zodOutputFormat(Contact) },
});
const contact = response.parsed_output; // typed, or null if parsing failed
if (!contact) throw new Error("No parsed output");
console.log(contact.plan); // "enterprise"The SDK validates the response against your schema and gives you parsed_output as a typed object, so there is no regex and no JSON.parse of text you hope is clean.
Strict tool arguments#
If the data is the arguments for something your code will do (book a flight, create a ticket), define it as a tool and set strict: true so the tool input is guaranteed to match its schema. More in tool use explained.
What can still go wrong#
Even with enforcement, check these:
stop_reasonismax_tokens. The reply was cut off. Raisemax_tokens.- A refusal. The model may decline a request; check
stop_reasonbefore you read content. - Valid shape, wrong values. The schema says
emailis a string, not that it is a real email. Validate business rules. - Too many optional fields. Smaller, tighter schemas give better extractions. Split big objects into steps.
if (response.stop_reason === "max_tokens") {
// retry with a higher max_tokens or a smaller schema
}Design schemas that models fill well#
- Use enums for closed sets instead of free text.
- Name fields clearly:
due_date_isobeatsdate. - Add a short
describe()to ambiguous fields. - Prefer required fields with an explicit
nulloption over a pile of optionals. - Keep it flat where you can.
Test it#
Structured output fixes the shape, not the accuracy. Build a small set of real inputs with known answers and score the extraction. See LLM evals explained. New to the API? Start with your first call in TypeScript.
Frequently asked questions
How do I get valid JSON from an LLM?
Use the API's structured output feature with a JSON schema so the response is constrained to match it. Prompt-only approaches usually work but occasionally return extra text or invalid JSON.
What is structured output?
A mode where you give the API a schema and the model's response is guaranteed to follow it, so you can parse it straight into a typed object.
Why does my LLM return JSON wrapped in markdown?
Models often add a code fence or a sentence before the JSON when only prompted. Schema-enforced output removes that, or you can strip fences and validate defensively.
Do I still need to validate structured output?
Yes. A valid shape does not mean valid values. Check ranges, enums and business rules, and handle the cases where the model refuses or the reply is cut off by max_tokens.
Prefer plain text? Read this page as Markdown.