# Claude API tutorial: your first call in TypeScript

> Make your first Claude API call in TypeScript: install the SDK, send a message, add a system prompt, keep a conversation, stream the reply and read token usage.

Source: https://devaiper.com/blog/claude-api-first-call-typescript
Published: 2026-10-08
Topics: Claude API, TypeScript, Tutorial

**Short answer:** Install @anthropic-ai/sdk, set ANTHROPIC_API_KEY, call client.messages.create with a model, max_tokens and a messages array, then read the text blocks in response.content. Add system for behavior, resend history for chat, and stream for long answers.

You can go from nothing to a working Claude call in about ten lines. This post builds it up step by step: a first call, a system prompt, a multi-turn chat, streaming and token counts. Everything is TypeScript and runs on Node 20 or newer.

## 1. Set up

```bash
mkdir claude-hello && cd claude-hello
npm init -y
npm pkg set type=module
npm i @anthropic-ai/sdk
npm i -D tsx typescript
```

Create an API key in the Claude Console and export it. The SDK reads `ANTHROPIC_API_KEY` automatically, so the key never appears in your code.

```bash
export ANTHROPIC_API_KEY="sk-ant-..."
```

## 2. Your first call

```ts
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic(); // reads ANTHROPIC_API_KEY

const response = await client.messages.create({
  model: "claude-opus-5-5",
  max_tokens: 1024,
  messages: [{ role: "user", content: "Explain what a closure is in one paragraph." }],
});

for (const block of response.content) {
  if (block.type === "text") console.log(block.text);
}
```

Run it with `npx tsx hello.ts`.

Three things to notice:

- **`model`** picks the model. Swap in a cheaper one (for example `claude-sonnet-5-5`) for high-volume work.
- **`max_tokens`** is a required ceiling on the reply length.
- **`response.content` is a list of blocks**, not a string. Text lives in blocks whose `type` is `"text"`, which is why you narrow before reading `.text`.

> **Diagram:** The shape of one Claude API call. Your code sends a request containing model, max_tokens, optional system prompt and a messages array. The API returns a response with a content array of blocks, a stop reason and token usage.
> Your request (model, max_tokens, system, messages) → Messages API → Response (content blocks, stop_reason, usage)
> One request in, one response out. Nothing is remembered between calls.

## 3. Add a system prompt

The system prompt sets behavior and rules. It is a **top-level field**, not a message.

```ts
const response = await client.messages.create({
  model: "claude-opus-5-5",
  max_tokens: 1024,
  system: "You are a senior TypeScript reviewer. Be direct. Point out bugs first.",
  messages: [{ role: "user", content: "const total = items.map(i => i.price).reduce((a, b) => a + b)" }],
});
```

## 4. Hold a conversation

The API is **stateless**. It does not remember the last call. To chat, you resend the whole history each time: your messages and the assistant's replies, alternating.

```ts
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();
const history: Anthropic.MessageParam[] = [];

async function ask(question: string): Promise<string> {
  history.push({ role: "user", content: question });

  const response = await client.messages.create({
    model: "claude-opus-5-5",
    max_tokens: 1024,
    messages: history,
  });

  const text = response.content
    .filter((b): b is Anthropic.TextBlock => b.type === "text")
    .map((b) => b.text)
    .join("");

  history.push({ role: "assistant", content: text });
  return text;
}

console.log(await ask("What is a closure?"));
console.log(await ask("Show me an example in a React hook."));
```

Because history grows each turn, so does the cost. Long chats need trimming or summarizing, which is a [context engineering](https://devaiper.com/blog/what-is-context-engineering) job.

## 5. Stream the reply

For long answers, stream so users see text immediately.

```ts
const stream = client.messages.stream({
  model: "claude-opus-5-5",
  max_tokens: 4096,
  messages: [{ role: "user", content: "Write a short guide to TypeScript generics." }],
});

stream.on("text", (delta) => process.stdout.write(delta));

const final = await stream.finalMessage();
console.log("\n\nstop reason:", final.stop_reason);
```

Use `finalMessage()` to get the complete message once it ends. No need to wrap events in a promise yourself.

## 6. Read the usage

Every response reports tokens, which is what you are billed on.

```ts
console.log(response.usage.input_tokens, response.usage.output_tokens);
```

Check `response.stop_reason` too: `end_turn` means it finished; `max_tokens` means you cut it off and should raise the limit.

## Common mistakes

| Mistake | Fix |
|---|---|
| Hardcoding the API key | Use the environment variable and keep it out of git |
| Reading `response.content[0].text` directly | Narrow on `block.type === "text"` first |
| Expecting memory between calls | Resend the history |
| `max_tokens` too small | Answers get cut off with `stop_reason: "max_tokens"` |
| Putting the system prompt in `messages` | Use the top-level `system` field |

## What to learn next

- [How to get reliable JSON from an LLM](https://devaiper.com/blog/how-to-get-json-from-an-llm)
- [Tool use (function calling) explained](https://devaiper.com/blog/tool-use-function-calling-explained)
- [Prompt caching: cut your bill](https://devaiper.com/blog/prompt-caching-explained)

## FAQ

### How do I call the Claude API from TypeScript?

Install @anthropic-ai/sdk, set the ANTHROPIC_API_KEY environment variable, create a client with new Anthropic(), and call client.messages.create with model, max_tokens and messages. The answer is in the text blocks of response.content.

### Does the Claude API remember my conversation?

No. The API is stateless. You send the full message history on every request, appending each assistant reply and the next user message.

### What is max_tokens?

The maximum number of tokens the model may generate in this reply. It is a ceiling on output length, not a target. Set it high enough that answers are not cut off.

### How do I stream a Claude response?

Use client.messages.stream, listen to the text event to print deltas as they arrive, and await stream.finalMessage() to get the complete message afterwards.

### Where do I put the system prompt?

In the top-level system field of the request, not as a message in the messages array.

