The gist
Up to now Claude has worked only with text, like an expert who reads but doesn't look. Multimodality gives it eyesight. Now you can show it a screenshot of an interface and ask "what's wrong here?", drop in a PDF invoice and ask it to "pull every amount into JSON", or snap a photo of a hand-drawn diagram and get code back. Claude Code "sees" images and "reads" documents directly, as a multimodal model, not through an OCR conversion step.
Key concepts
- Vision: Claude analyzes images through the API (application programming interface): base64, URL or the Files API
- PDF support: the Read tool in Claude Code handles PDFs automatically (with page ranges)
- Supported formats: JPEG, PNG, GIF, WebP for images; PDF for documents
- The token formula (a token is a unit of text for AI): an image is cut into 28x28 px blocks, and each block = 1 token:
⌈width/28⌉ × ⌈height/28⌉(from the official documentation) - Limits (as of October 2026): up to 10 MB per file directly through the API (5 MB on Amazon Bedrock and Google Cloud, 10 MB on claude.ai), up to 8000x8000 px, up to 600 images in one request (100 for models with a 200K window, such as Haiku 4.5)
- Native resolution: up to 1568 px on the long side; 2576 px for Claude 4.7 models and newer (Opus 5.5, Sonnet 5.5, Fable 5.1)
Theory
How Claude sees images
Claude doesn't receive an image as a grid of pixels. It receives it as encoded data and processes it in its multimodal space. That means it can:
- Read text in screenshots, photos of documents, handwriting
- Analyze diagrams, charts, schematics
- Describe a UI: "the Send button is at the bottom right, the form has 3 fields"
- Find mistakes: compare a design mockup with the actual build
- Pull structured data out of tables and forms
What Claude CANNOT do with images (from Anthropic's documentation):
- Generate images: Claude only analyzes; it doesn't create or edit them
- Identify people: Claude doesn't name people in photos (an Acceptable Use Policy restriction)
- Judge spatial relationships precisely: it can get analog clocks or chess positions wrong
- Guarantee an exact count of objects: it gives approximate numbers
- Detect AI-generated images: it can't tell a fake from a real photo
- Analyze video: only still frames; for animated GIFs, only the first frame
Multimodality in Claude Code
In Claude Code, you work with images through the Read tool, the same tool that reads text files:
Read ~/screenshots/ui-bug.png and find problems with the interface
Claude Code detects the file type automatically. For images it takes in the content visually (Claude is a multimodal model). PNG, JPG, GIF and WebP are supported.
For PDFs there's a pages parameter for reading specific pages:
Read ~/docs/contract.pdf pages 1-5 and pull out the key points
Three ways to pass an image through the API
Option 1: URL (if the image is publicly available)
import anthropic
client = anthropic.Anthropic()
message = client.messages.create(
model="claude-opus-5-5",
max_tokens=1024,
messages=[
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "url",
"url": "https://example.com/screenshot.png",
},
},
{
"type": "text",
"text": "What does this screenshot show? Describe the UI elements."
}
],
}
],
)
print("".join(b.text for b in message.content if b.type == "text"))Option 2: Base64 (for local files)
import anthropic
import base64
# Read the file and encode it as base64
with open("invoice.png", "rb") as f:
image_data = base64.standard_b64encode(f.read()).decode("utf-8")
client = anthropic.Anthropic()
message = client.messages.create(
model="claude-opus-5-5",
max_tokens=1024,
messages=[
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/png", # image/jpeg, image/gif, image/webp
"data": image_data,
},
},
{
"type": "text",
"text": "Extract from this invoice: the number, date, total amount and recipient. Return JSON."
}
],
}
],
)
print("".join(b.text for b in message.content if b.type == "text"))Option 3: Files API (upload once, reuse many times)
For images you use repeatedly, or in long conversations, the Files API lets you upload a file once and refer to it by file_id:
import anthropic
client = anthropic.Anthropic()
# Upload the file once
with open("image.jpg", "rb") as f:
file_upload = client.files.upload(file=("image.jpg", f, "image/jpeg"))
# Use the file_id (no need to send base64 every time)
message = client.messages.create(
model="claude-opus-5-5",
max_tokens=1024,
messages=[
{
"role": "user",
"content": [
{
"type": "image",
"source": {"type": "file", "file_id": file_upload.id},
},
{"type": "text", "text": "Describe this image."},
],
}
],
)Why the Files API? (It used to require a beta header; according to the current documentation, it no longer does.) In multi-turn conversations, every request resends the whole history. If images are encoded as base64, they're sent in full on every turn. With the Files API, only the file_id is sent, so the request doesn't keep growing.
Several images in one request
Claude can work with several images at once, for example to compare a design mockup with a screenshot of the build.
Limits (from Anthropic's documentation):
- Up to 20 images per message on claude.ai
- Up to 100 images per request through the API (models with a 200K context, currently Haiku 4.5)
- Up to 600 images per request through the API (other models)
- With more than 20 images in one request, a stricter size limit applies to each image (the documentation says no more than 2000 px per side)
Best practice: with several images, label each one: "Image 1:", "Image 2:":
message = client.messages.create(
model="claude-opus-5-5",
max_tokens=1024,
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Image 1:"},
{
"type": "image",
"source": {"type": "base64", "media_type": "image/png", "data": design_b64},
},
{"type": "text", "text": "Image 2:"},
{
"type": "image",
"source": {"type": "base64", "media_type": "image/png", "data": screenshot_b64},
},
{
"type": "text",
"text": "Find the differences between the mockup (Image 1) and the build (Image 2)."
}
],
}
],
)Order matters: according to Anthropic's documentation, images work better placed before the text question, the same as with long documents.
PDFs: the Read tool in Claude Code
In interactive mode, Claude Code handles PDFs through the built-in Read tool. Just give it the file path:
Read ~/documents/contract.pdf and pull out every payment deadline
Claude Code automatically:
- Opens the PDF
- Extracts the text, tables and structure
- Works with the content
For large PDFs (more than 10 pages), always give a page range:
Read ~/docs/report.pdf pages 1-5
The maximum is 20 pages per request. Without a range, reading a large PDF may fail with an error.
Through the API: a PDF as base64:
import anthropic
import base64
with open("report.pdf", "rb") as f:
pdf_data = base64.standard_b64encode(f.read()).decode("utf-8")
client = anthropic.Anthropic()
message = client.messages.create(
model="claude-opus-5-5",
max_tokens=2048,
messages=[
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": pdf_data,
},
},
{
"type": "text",
"text": "Pull every numeric metric out of this report as a Markdown table."
}
],
}
],
)Cost: how tokens are calculated
The official formula from Anthropic's documentation: Claude looks at an image in 28x28 pixel blocks, and each block is one token.
tokens = ⌈width / 28⌉ × ⌈height / 28⌉Width and height are in pixels, after resizing. (Older versions of this course used the formula width * height / 750; it's out of date.)
Maximum native resolution (as of October 2026):
- For Claude 4.7 models and newer: 2576 px on the long side (up to 4784 tokens per image)
- For other models: 1568 px on the long side (up to 1568 tokens per image)
If an image is larger, Claude automatically scales it down and keeps the proportions.
Examples at the standard resolution level (for example Haiku 4.5, $1 per 1M input tokens) and at the high level (Opus 5.5, $4 per 1M). Prices as of October 2026; for current ones, see What's current.
| Image size | Tokens (standard level) | Cost (Haiku 4.5, $1/1M) |
|---|---|---|
| 200x200 px | 64 | ~$0.00006 |
| 1000x1000 px | 1296 | ~$0.0013 |
| 1920x1080 px | 1560 (scaled down) | ~$0.0016 |
| Image size | Tokens (high resolution) | Cost (Opus 5.5, $4/1M) |
|---|---|---|
| 200x200 px | 64 | ~$0.00026 |
| 1000x1000 px | 1296 | ~$0.0052 |
| 1920x1080 px | 2691 (high resolution) | ~$0.011 |
Claude 4.7 models and newer support high resolution. That's up to 3x more tokens per image, but they see fine detail better. If you don't need high precision, shrink the image before sending it.
Image quality tips (from Anthropic's documentation)
- Format: JPEG, PNG, GIF or WebP. Animations aren't supported; only the first frame is used
- Sharpness: the image should be sharp, not blurry or pixelated
- Text: if the image has important text, make sure it's readable and not too small. Don't crop out visual context just to make the text bigger
- Compression: lossy JPEG/WebP compression reduces request size and latency, but it can create artifacts. Check that text is still readable after compression
- Resizing: keep in mind that the image may be scaled down automatically, which can make small text unreadable. It's better to resize it yourself ahead of time
Practical use cases
Analyzing UI screenshots:
Here's a screenshot of our checkout page [image]. Find UX problems: what might confuse the user, where there isn't enough contrast, which elements are placed in a confusing way.
Extracting data from documents:
Here's a photo of a restaurant receipt [image]. Extract: the restaurant name, the date, each item with its price, the total, and the tip if there is one. Return structured JSON.
Turning hand-drawn diagrams into code:
Here's a photo of a hand-drawn database diagram [image]. Convert it into SQL CREATE TABLE statements for PostgreSQL.
Two-pass visual check (the Build → Screenshot → Review → Fix pattern):
A powerful pattern for UI work: first you write the code, then Claude checks the visual result:
# Step 1: Claude generates HTML/CSS
# Step 2: Take a screenshot of the page (with Playwright, Puppeteer, or Computer Use)
# Step 3: Send the screenshot to Claude and ask "what's wrong?"
# Step 4: Claude finds visual bugs and fixes the code
message = client.messages.create(
model="claude-opus-5-5",
max_tokens=2048,
messages=[{
"role": "user",
"content": [
{"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": screenshot_b64}},
{"type": "text", "text": "This is a screenshot of a page I just generated. "
"Find visual problems: misalignment, cut-off text, "
"contrast issues, wrong spacing. Suggest CSS fixes."}
]
}]
)Automated design QA (comparing Figma with the build):
# Compare each component from Figma with a real screenshot
for component_name, figma_img, screenshot_img in components:
result = client.messages.create(
model="claude-opus-5-5",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Image 1 (Figma mockup):"},
{"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": figma_img}},
{"type": "text", "text": "Image 2 (build):"},
{"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": screenshot_img}},
{"type": "text", "text": f"Component: {component_name}. Rate the match from 0 to 10 and describe the differences."}
]
}]
)Practice
Assignment 1: Invoice analyzer
- Find or create a PNG image of an invoice. You can take a screenshot of any template online
- Write a Python script that reads the file, encodes it as base64 and sends it to Claude with a prompt (a prompt is your request to the AI):
"Extract from this invoice: the document number, date, recipient, the list of line items (name, quantity, price) and the total amount. Return JSON." - Run the script and check that the extracted data is correct
- Add error handling: when the file isn't found, when Claude couldn't extract the data
- Bonus: process several invoices in a loop and save the results to
invoices.json
Assignment 2: A two-pass visual check in Claude Code
- Ask Claude Code to build an HTML page:
"Create a landing page for a fitness app with a hero section, pricing and a CTA button" - Open the result in your browser and take a screenshot (or use the Claude Code Read tool to view it)
- Show the screenshot to Claude:
"Read /tmp/screenshot.png and find visual problems with this page" - Claude will suggest fixes; apply them
- Repeat the loop: screenshot → review → fix, until you're happy with the result
Goal: understand how to pass images through the API and pull structured data out of visual sources. Get comfortable with the two-pass visual check pattern.
Tools and resources
- anthropic:
pip install anthropic, the Python SDK with vision support - Pillow:
pip install Pillow, resizes images before sending (saves tokens) - base64: a built-in Python module for encoding files
- Claude Code Read tool: built-in support for images (PNG, JPG, GIF, WebP) and PDFs
- Files API: upload and reuse images through
file_id - Vision documentation: platform.claude.com/docs/en/build-with-claude/vision
- Multimodal cookbook: platform.claude.com/cookbook/multimodal-getting-started-with-vision
Key takeaways
You can pass an image to the API in three ways: base64, URL, or the Files API (file_id). The type goes in the
media_typefield. Claude handles JPEG, PNG, GIF, WebP and PDF.
In Claude Code, images and PDFs open through the Read tool like regular files; no conversion needed. For large PDFs (more than 10 pages), give a page range.
Cost =
⌈width/28⌉ × ⌈height/28⌉tokens. The maximum is 1568px on the long side (2576px for Claude 4.7 models and newer). Shrink images to the smallest size that works before sending.
Claude does NOT generate images; it only analyzes them. It doesn't identify people in photos. It can make mistakes on tiny, blurry or rotated images.
The two-pass visual check pattern (Build → Screenshot → Review → Fix) is a powerful tool for UI work, with Claude doing the visual verification.
Next lesson
→ AI Image Generation Pipeline: how to make images when Claude doesn't draw them. After that: What are Skills.
The mark stays in this browser only and is never sent anywhere. My progress