Skip to content

Lesson 12: Getting started with Claude Code and the LiteLLM proxy

Over the first eleven lessons, we built a strong foundation in how agents work - reasoning, tools, memory, planning, multi-agent systems, RAG, evaluation, safety, and production operations. All of those concepts are platform-agnostic; they apply whether you build on Claude, GPT, Gemini, or an open-source model.

Now we shift from theory to practice. From here on, the course is hands-on, and everything you build uses one concrete stack:

  • Claude (Claude Enterprise) - the chat, documents, and projects surface you already use in the browser and desktop app.
  • Claude Code - the terminal-based coding agent that is already installed and working on your machine. This is your daily driver.
  • The Anthropic API, reached through your company’s LiteLLM proxy - the programmatic entry point for everything you write in code. One endpoint, one bill, Claude and OpenAI models behind it.
  • The Claude Agent SDK - the framework you use to build your own agents, covered in the next lesson.

This lesson maps the concepts from earlier lessons onto these tools, explains how the pieces fit together, and shows you how to make your first API call through the proxy. By the end you will know which tool to reach for, how to get and configure your API key, and how to run a Claude model from the terminal, from curl, and from Python.

⚠️ Safety first: Before you build anything with these tools, revisit Lesson 10: Guardrails and safety. Your company defines its own AI principles and guardrails on top of the general practices in that lesson - they are documented internally and take precedence over anything written here. Review them before you write your first agent.

ELI5: Think of the stack like a professional kitchen

Section titled “ELI5: Think of the stack like a professional kitchen”

Imagine you are cooking in a busy restaurant kitchen rather than at home. You need ingredients (raw food), a set of knives and pans (the tools you actually work with), a pass where orders come in and plated food goes out (a single controlled point everything flows through), and a recipe book that turns individual techniques into full dishes.

Our Claude stack works the same way:

  • Claude models are your ingredients - the intelligence that powers everything.
  • Claude Code is your knife roll - the tool you pick up every day to get real work done.
  • The LiteLLM proxy is the pass - every programmatic order flows through it, so the kitchen can track what is cooked, who ordered it, and how much it costs.
  • The Claude Agent SDK is your recipe book - it packages the techniques from Claude Code into repeatable, buildable agents.
  • Claude Enterprise chat is the dining room - where people who are not in the kitchen still get served.

You can use these pieces independently or together. Not every task needs every tool - sometimes you just ask Claude in the browser; sometimes you write code against the API.

Key takeaway: The course uses a single, coherent stack - Claude Code for interactive work, the Anthropic API through your company’s LiteLLM proxy for programmatic work, and the Claude Agent SDK to build agents. Knowing which piece does what lets you pick the right tool for each job.


Before we look at the pieces, it helps to understand why your programmatic calls do not go straight to Anthropic. Instead, they pass through a LiteLLM proxy that your company runs.

LiteLLM is an open-source proxy that speaks the Anthropic and OpenAI API formats and forwards requests to the real providers. Putting one in front of the model providers gives a company several things that matter at scale:

  • A single endpoint and a single bill. Every team calls one URL. Finance sees one invoice instead of a separate contract with each model provider.
  • Per-team budgets and rate limits. The proxy issues keys scoped to a team or project, each with its own spending cap and request-rate ceiling. One runaway script cannot burn the whole company’s budget.
  • Central logging and audit. Because every request flows through one place, the platform team can log usage, trace incidents, and answer “who called which model with what” - which matters for both cost control and safety review.
  • Model allow-lists. The proxy decides which models are exposed. That is how a company enforces “we only use approved models” without trusting every developer to remember the policy.
  • Painless A/B of providers. Because the proxy exposes Anthropic and OpenAI models behind the same API shape, you can compare a Claude model against a GPT model by changing a single string - no rewrite, no second SDK.

The practical upshot for you: you write ordinary Anthropic SDK code, point two environment variables at the proxy, and everything else just works. We will do exactly that below. (Not every workload goes through the proxy - some reach Claude directly via AWS Bedrock instead; the proxy is the default path for course work, so follow the internal documentation to know when Bedrock applies to you.)


Here is a map of the stack and how the parts relate:

+------------------------------------------------------------------+
| You, building things |
+------------------------------------------------------------------+
| |
| +--------------------+ +----------------------------------+ |
| | Claude (chat) | | Claude Code | |
| | Claude Enterprise | | (terminal agent) | |
| | | | | |
| | - Ask questions | | - Read & change code | |
| | - Documents | | - Run commands | |
| | - Projects | | - MCP, skills, subagents | |
| +--------------------+ +----------------------------------+ |
| | |
| v |
| +----------------------------------------------------+ |
| | Claude Agent SDK | |
| | (the agent loop from Claude Code, as a library) | |
| +----------------------------------------------------+ |
| | |
| v |
| +----------------------------------------------------+ |
| | Anthropic API (client.messages.create) | |
| +----------------------------------------------------+ |
| | |
| v |
| +----------------------------------------------------+ |
| | Your company's LiteLLM proxy (one endpoint) | |
| | budgets - rate limits - logging - allow-list | |
| +----------------------------------------------------+ |
| | |
| v |
| +----------------------------------------------------+ |
| | Claude models (Opus 4.8 - Sonnet 5 - Haiku 4.5) | |
| | + recent OpenAI GPT-series models | |
| +----------------------------------------------------+ |
+------------------------------------------------------------------+
Claude Stack Explorer
Click any piece to learn more. Try the Decision Helper.
What you use directly
💬
Claude (chat)
Claude Enterprise
⌨️
Claude Code
Terminal agent
▼ ▼
What you build with
🧰
Claude Agent SDK
Build agents
🔌
Anthropic API
Raw messages call
▼ ▼
The gateway and the models
🛂
LiteLLM proxy
One endpoint - budgets, limits, logging, allow-list
🧠
Opus 4.8
Most capable
Sonnet 5
Balanced
💨
Haiku 4.5
Fastest, cheapest

Here is the same map as a table, with the question each piece answers:

PieceWhat it isReach for it when
Claude (chat)Claude Enterprise chat, documents, and projectsYou want to ask a question, draft or review a document, or think out loud - no code
Claude CodeTerminal coding agent, already installedYou are working in a repository: explaining, changing, running, automating
Anthropic API via the proxyThe raw messages endpoint through your company gatewayYou want full, framework-free control from your own program
Claude Agent SDKThe Claude Code agent loop, as a libraryYou are building an agent as a product or automation

The rule of thumb: use chat to think, Claude Code to work, the API for control, and the SDK to build. All four talk to the same models through the same proxy.


Claude Code and Claude chat are already set up for you. To call models from your own code, you need an API key for the proxy.

Keys come from your internal developer portal. The exact steps, the portal’s location, and the proxy endpoint are documented internally - follow that documentation rather than anything hard-coded here (it can change, and it is not public). In general terms you will:

  1. Sign in to the internal developer portal with your company identity.
  2. Request or create a key for your team or project. The key is typically scoped to a budget and a rate limit, and to the set of models your team is allowed to use.
  3. Copy the key once. Most portals show the secret only at creation time.

An API key is a credential that can spend real money and reach real models. Treat it like a password:

  • Never commit it. Keep keys out of source control entirely. Add .env files to .gitignore; do not paste keys into code, notebooks, or chat messages.
  • Use environment variables, or a secrets manager for anything shared or deployed. The Anthropic SDKs read the key from the environment automatically (below), so you never need it in a source file.
  • Scope narrowly. Use a key tied to your team and its budget, not a broad shared one. That keeps usage attributable and blast radius small.
  • Rotate on exposure. If a key ever lands somewhere it should not - a commit, a log, a screenshot - revoke and reissue it through the portal. Rotate periodically as a matter of routine, too.

If a key leaks, the proxy’s per-key budget and logging limit the damage and make it visible - another reason the gateway exists - but rotating fast is still on you.


Configuring your environment and making your first calls

Section titled “Configuring your environment and making your first calls”

The Anthropic SDKs and the curl examples below read two environment variables: ANTHROPIC_BASE_URL (where to send requests) and ANTHROPIC_API_KEY (your credential). Point the first at the proxy and the second at your key:

Terminal window
export ANTHROPIC_BASE_URL="https://<your-company-llm-proxy>" # endpoint from the internal docs
export ANTHROPIC_API_KEY="sk-..." # key from the internal developer portal

Because the SDKs pick up ANTHROPIC_BASE_URL automatically, every piece of Anthropic SDK code in this course works through the proxy with no proxy-specific changes. You write ordinary Anthropic code; the two variables route it through your company’s gateway. That is the whole trick.

The quickest way to confirm your key and endpoint work is a raw HTTP request:

Terminal window
curl "$ANTHROPIC_BASE_URL/v1/messages" \
-H "content-type: application/json" \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-d '{
"model": "aws/claude-4-8-opus",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Say hello!"}]
}'

A JSON response with a content array means everything is wired up correctly. An authentication error usually means the key or the base URL is wrong; a model error usually means the alias is not one your proxy exposes (more on that below).

For real work you will use the SDK. Install it and make the same call:

Terminal window
uv add anthropic
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY and ANTHROPIC_BASE_URL from the environment
response = client.messages.create(
model="aws/claude-4-8-opus",
max_tokens=1024,
messages=[{"role": "user", "content": "Say hello!"}],
)
for block in response.content:
if block.type == "text":
print(block.text)

Note that response.content is a list of blocks, not a single string. A reply can contain text blocks, tool-use blocks, and more - iterating over them and checking block.type is the pattern you will use throughout the course. We build directly on this loop in the next lesson.

The model aliases available through the proxy come from the proxy’s own configuration, not from Anthropic directly. You can list what your proxy exposes:

Terminal window
curl "$ANTHROPIC_BASE_URL/v1/models" -H "x-api-key: $ANTHROPIC_API_KEY"

The names it returns are set by the platform team and differ from Anthropic’s official IDs - and the exact spellings differ a bit in every company’s environment. This course’s examples use the identifiers as they appear on our proxy (for example aws/claude-4-8-opus); if a model string is rejected, check /v1/models for the exact string your proxy expects and use that. The code shapes in this course stay the same; only the model string changes.


Your proxy exposes three Claude models. They are the same three you will use for every code example in the course (remember: the exact identifier spellings are proxy-specific and vary per company):

ModelIDBest forCharacteristics
Claude Opus 4.8aws/claude-4-8-opusComplex reasoning, hard planning, quality-critical workMost capable, higher latency and relative cost
Claude Sonnet 5aws/claude-5-sonnetGeneral agent work - tool use, RAG, summarization, chatBalanced speed and intelligence; the workhorse
Claude Haiku 4.5aws/claude-4-5-haikuClassification, routing, extraction, high-volume simple stepsFastest and cheapest

The proxy also exposes recent OpenAI GPT-series models behind the same gateway - currently three variants from the GPT-5.6 family:

ModelID
GPT-5.6 (Luna)azure/gpt-5-6-luna
GPT-5.6 (Sol)azure/gpt-5-6-sol
GPT-5.6 (Terra)azure/gpt-5-6-terra

That makes it easy to compare a Claude model against a GPT model for a given task by swapping the model string - useful for evaluation (Lesson 9) and for the kind of A/B testing the gateway is built to support. This course teaches on the Claude models, but the mechanism is provider-neutral.

Claude Opus 4.8 and Sonnet 5 support context windows up to 1M tokens; Haiku 4.5 supports 200K tokens. That is enough to put entire codebases, long document sets, or lengthy conversations in a single request - though larger contexts cost more and can be slower, so use only what the task needs.

Think of model selection like choosing the right vehicle for a trip:

  • Haiku 4.5 is a bicycle - fast, cheap, ideal for short trips (classification, routing, simple extraction).
  • Sonnet 5 is a car - versatile, good for most journeys (general agent tasks, tool use, RAG, conversation). Make it your default.
  • Opus 4.8 is a truck - powerful, handles heavy loads (complex multi-step reasoning, long documents, difficult planning).

Most production agents use more than one tier through model routing (Lesson 11): Haiku for the simple steps, Sonnet for the core logic, and Opus only when the task genuinely requires it. Because relative cost rises steeply from Haiku to Opus, routing well is one of the biggest levers you have on spend. (Absolute prices depend on your company’s proxy arrangement - think in relative tiers, and check the internal docs for actuals.)


Claude Code is where you will spend most of your hands-on time. It is an agent that lives in your terminal, sees your project, and can read code, make changes, and run commands on your behalf - the ReAct-style loop from Lesson 4, made concrete. Because it is already installed, you can try these right now in any repository.

A quick tour of what it is good at:

  • Explaining code. “Walk me through how authentication works in this service.” Claude Code reads the relevant files and explains them, with references you can click.
  • Making changes. “Add input validation to the signup handler and a test for it.” It edits the files, runs the test, and shows you the diff. You stay in control - it asks before doing anything irreversible, the permission model from Lesson 10 in practice.
  • Running commands. “Why is the build failing?” It runs the build, reads the error, and proposes a fix.

Three features turn Claude Code from a helper into a tailored teammate. Each has its own lesson later in the course:

  • Project context with CLAUDE.md. A CLAUDE.md file at the root of your repo gives Claude Code standing instructions - conventions, architecture notes, commands, do’s and don’ts - that it reads on every session. This is the single highest-leverage way to make it fit your codebase. We cover it in Lesson 15: AGENTS.md and project context.
  • Extending with MCP. The Model Context Protocol lets Claude Code connect to external tools and data sources - issue trackers, databases, internal services - through a standard interface. We introduce it in Lesson 14: Agent protocols - MCP and A2A and go deep in Lesson 16: MCP deep dive.
  • Packaging know-how as skills. Skills bundle reusable instructions and resources that Claude Code loads when they are relevant - your team’s playbook for a recurring task, made executable. We cover them in Lesson 17: Agent skills.

For the full feature set - slash commands, subagents, hooks, IDE integrations - see the Claude Code documentation. Everything Claude Code does programmatically flows through the same proxy your API keys point at, so it stays inside your company’s budgets, limits, and audit trail.


Where each lesson concept maps to the Claude stack

Section titled “Where each lesson concept maps to the Claude stack”

Here is a reference connecting the concepts from earlier lessons to the tools you will use from here on:

LessonConceptClaude stack
2 - How agents thinkLLM reasoningClaude models (Opus 4.8, Sonnet 5, Haiku 4.5)
3 - ToolsFunction callingAnthropic API tools; MCP tools in Claude Code/SDK
4 - Design patternsOrchestrationThe agent loop; Claude Agent SDK subagents
5 - MemorySession state, long-term memoryConversation history; CLAUDE.md; SDK sessions
6 - PlanningMulti-step reasoningOpus 4.8 for complex planning
7 - Multi-agentAgent coordinationClaude Agent SDK subagents; A2A (Lesson 14)
8 - RAGKnowledge retrievalMCP data sources; your own retrieval + the API
9 - EvaluationTesting agentsA/B of Claude and GPT models through the proxy
10 - SafetyGuardrailsClaude Code permissions; company guardrails via the proxy
11 - ProductionDeployment, CI/CDClaude Agent SDK; proxy budgets, limits, and logging

Beyond this course’s core stack, these are the standard tool choices for production agent work at the company; later lessons introduce each where it becomes relevant. None of them replace what you are about to build - they sit alongside it once an agent moves from prototype to production.

ConcernToolWhere in this course
Language & packagingPython 3.14 + uvUsed throughout
HTTP servicesFastAPI + uvicornLesson 11
Structured LLM outputInstructor (+ Pydantic)Lesson 13
Claude-native agentsClaude Agent SDKLesson 13
MCP servers (Python)FastMCPLesson 14 and 16
OrchestrationLangGraphLesson 18
Vector store / vectorized memoryQdrantLesson 5 and 8
Tracing & evaluationLangfuseLesson 9 and 11

  1. The course runs on one coherent stack. Claude for chat, Claude Code for interactive work, the Anthropic API through your company’s LiteLLM proxy for programmatic work, and the Claude Agent SDK to build agents. All four use the same models through the same gateway.

  2. The proxy is why things stay sane at scale. One endpoint and one bill, per-team budgets and rate limits, central logging for audit, a model allow-list, and painless A/B of Anthropic and OpenAI models. You get all of it by setting two environment variables.

  3. Configuration is two variables. Export ANTHROPIC_BASE_URL and ANTHROPIC_API_KEY, and every Anthropic SDK example in this course runs through the proxy unchanged.

  4. Keys are credentials. They come from your internal developer portal. Never commit them, keep them in environment variables or a secrets manager, scope them narrowly, and rotate on exposure.

  5. Pick the right model tier. Haiku 4.5 for simple, high-volume steps; Sonnet 5 as your default; Opus 4.8 for the hard cases. Route across tiers to control cost. Model aliases come from the proxy’s /v1/models and may differ from Anthropic’s official IDs.

  6. Claude Code is your daily driver. Explain, change, and run code interactively, then tailor it with CLAUDE.md (Lesson 15), MCP (Lessons 14 and 16), and skills (Lesson 17).



Previous lesson: From Prototype to Production

Next lesson: Building Your First Agent