Library · Connections: APIs, MCP and running 24/7

Multimodal Claude (working with more than text): images, PDFs, audio

Builder50 minUpdated: October 2026
26 of 105 in the library

Time: about 20 min reading + 30 min practice


The gist

Up to now Claude has worked only with text, like an expert who reads but doesn't look. Multimodality gives it eyesight. Now you can show it a screenshot of an interface and ask "what's wrong here?", drop in a PDF invoice and ask it to "pull every amount into JSON", or snap a photo of a hand-drawn diagram and get code back. Claude Code "sees" images and "reads" documents directly, as a multimodal model, not through an OCR conversion step.


Key concepts

  • Vision: Claude analyzes images through the API (application programming interface): base64, URL or the Files API
  • PDF support: the Read tool in Claude Code handles PDFs automatically (with page ranges)
  • Supported formats: JPEG, PNG, GIF, WebP for images; PDF for documents
  • The token formula (a token is a unit of text for AI): an image is cut into 28x28 px blocks, and each block = 1 token: ⌈width/28⌉ × ⌈height/28⌉ (from the official documentation)
  • Limits (as of October 2026): up to 10 MB per file directly through the API (5 MB on Amazon Bedrock and Google Cloud, 10 MB on claude.ai), up to 8000x8000 px, up to 600 images in one request (100 for models with a 200K window, such as Haiku 4.5)
  • Native resolution: up to 1568 px on the long side; 2576 px for Claude 4.7 models and newer (Opus 5.5, Sonnet 5.5, Fable 5.1)

Theory

How Claude sees images

🎨 Picture this: multimodality is giving eyesight to a blind expert. Before, they knew everything, but only by touch. Now you show them the screen, and they instantly see what's wrong with the interface, which numbers are in the PDF, what the diagram looks like.

Claude doesn't receive an image as a grid of pixels. It receives it as encoded data and processes it in its multimodal space. That means it can:

  • Read text in screenshots, photos of documents, handwriting
  • Analyze diagrams, charts, schematics
  • Describe a UI: "the Send button is at the bottom right, the form has 3 fields"
  • Find mistakes: compare a design mockup with the actual build
  • Pull structured data out of tables and forms

What Claude CANNOT do with images (from Anthropic's documentation):

  • Generate images: Claude only analyzes; it doesn't create or edit them
  • Identify people: Claude doesn't name people in photos (an Acceptable Use Policy restriction)
  • Judge spatial relationships precisely: it can get analog clocks or chess positions wrong
  • Guarantee an exact count of objects: it gives approximate numbers
  • Detect AI-generated images: it can't tell a fake from a real photo
  • Analyze video: only still frames; for animated GIFs, only the first frame

Multimodality in Claude Code

In Claude Code, you work with images through the Read tool, the same tool that reads text files:

Type this into the chat
Read ~/screenshots/ui-bug.png and find problems with the interface

Claude Code detects the file type automatically. For images it takes in the content visually (Claude is a multimodal model). PNG, JPG, GIF and WebP are supported.

For PDFs there's a pages parameter for reading specific pages:

Type this into the chat
Read ~/docs/contract.pdf pages 1-5 and pull out the key points

🎨 Picture this: the Read tool in Claude Code is a universal scanner. You put a document on it (text, image, PDF), and Claude sees the content directly. Nothing needs converting or re-encoding.


Three ways to pass an image through the API

Option 1: URL (if the image is publicly available)

python
import anthropic

client = anthropic.Anthropic()

message = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=1024,
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "image",
                    "source": {
                        "type": "url",
                        "url": "https://example.com/screenshot.png",
                    },
                },
                {
                    "type": "text",
                    "text": "What does this screenshot show? Describe the UI elements."
                }
            ],
        }
    ],
)
print("".join(b.text for b in message.content if b.type == "text"))

Option 2: Base64 (for local files)

python
import anthropic
import base64

# Read the file and encode it as base64
with open("invoice.png", "rb") as f:
    image_data = base64.standard_b64encode(f.read()).decode("utf-8")

client = anthropic.Anthropic()

message = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=1024,
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "image",
                    "source": {
                        "type": "base64",
                        "media_type": "image/png",  # image/jpeg, image/gif, image/webp
                        "data": image_data,
                    },
                },
                {
                    "type": "text",
                    "text": "Extract from this invoice: the number, date, total amount and recipient. Return JSON."
                }
            ],
        }
    ],
)
print("".join(b.text for b in message.content if b.type == "text"))

Option 3: Files API (upload once, reuse many times)

For images you use repeatedly, or in long conversations, the Files API lets you upload a file once and refer to it by file_id:

python
import anthropic

client = anthropic.Anthropic()

# Upload the file once
with open("image.jpg", "rb") as f:
    file_upload = client.files.upload(file=("image.jpg", f, "image/jpeg"))

# Use the file_id (no need to send base64 every time)
message = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=1024,
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "image",
                    "source": {"type": "file", "file_id": file_upload.id},
                },
                {"type": "text", "text": "Describe this image."},
            ],
        }
    ],
)

Why the Files API? (It used to require a beta header; according to the current documentation, it no longer does.) In multi-turn conversations, every request resends the whole history. If images are encoded as base64, they're sent in full on every turn. With the Files API, only the file_id is sent, so the request doesn't keep growing.

🎨 Picture this: base64 is printing out the full blueprint every time and carrying it to the meeting. The Files API is saying "look at blueprint #7, the one we already approved." One label instead of a heavy stack of paper.


Several images in one request

Claude can work with several images at once, for example to compare a design mockup with a screenshot of the build.

Limits (from Anthropic's documentation):

  • Up to 20 images per message on claude.ai
  • Up to 100 images per request through the API (models with a 200K context, currently Haiku 4.5)
  • Up to 600 images per request through the API (other models)
  • With more than 20 images in one request, a stricter size limit applies to each image (the documentation says no more than 2000 px per side)

🎨 Picture this: several images in one request is like showing an architect two floor plans at once. "Here's plan 1 (what we designed), here's plan 2 (what got built). Find the differences." Without labels, the architect will mix up which is which.

Best practice: with several images, label each one: "Image 1:", "Image 2:":

python
message = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=1024,
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Image 1:"},
                {
                    "type": "image",
                    "source": {"type": "base64", "media_type": "image/png", "data": design_b64},
                },
                {"type": "text", "text": "Image 2:"},
                {
                    "type": "image",
                    "source": {"type": "base64", "media_type": "image/png", "data": screenshot_b64},
                },
                {
                    "type": "text",
                    "text": "Find the differences between the mockup (Image 1) and the build (Image 2)."
                }
            ],
        }
    ],
)

Order matters: according to Anthropic's documentation, images work better placed before the text question, the same as with long documents.


PDFs: the Read tool in Claude Code

In interactive mode, Claude Code handles PDFs through the built-in Read tool. Just give it the file path:

Type this into the chat
Read ~/documents/contract.pdf and pull out every payment deadline

Claude Code automatically:

  1. Opens the PDF
  2. Extracts the text, tables and structure
  3. Works with the content

For large PDFs (more than 10 pages), always give a page range:

Type this into the chat
Read ~/docs/report.pdf pages 1-5

The maximum is 20 pages per request. Without a range, reading a large PDF may fail with an error.

🎨 Picture this: the Read tool for PDFs is an accountant who can read any document without a scanner. You put a contract on the desk, and they immediately see every amount, date and condition. Nothing needs converting.

Through the API: a PDF as base64:

python
import anthropic
import base64

with open("report.pdf", "rb") as f:
    pdf_data = base64.standard_b64encode(f.read()).decode("utf-8")

client = anthropic.Anthropic()

message = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=2048,
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "document",
                    "source": {
                        "type": "base64",
                        "media_type": "application/pdf",
                        "data": pdf_data,
                    },
                },
                {
                    "type": "text",
                    "text": "Pull every numeric metric out of this report as a Markdown table."
                }
            ],
        }
    ],
)

Cost: how tokens are calculated

The official formula from Anthropic's documentation: Claude looks at an image in 28x28 pixel blocks, and each block is one token.

Code
tokens = ⌈width / 28⌉ × ⌈height / 28⌉

Width and height are in pixels, after resizing. (Older versions of this course used the formula width * height / 750; it's out of date.)

Maximum native resolution (as of October 2026):

  • For Claude 4.7 models and newer: 2576 px on the long side (up to 4784 tokens per image)
  • For other models: 1568 px on the long side (up to 1568 tokens per image)

If an image is larger, Claude automatically scales it down and keeps the proportions.

Examples at the standard resolution level (for example Haiku 4.5, $1 per 1M input tokens) and at the high level (Opus 5.5, $4 per 1M). Prices as of October 2026; for current ones, see What's current.

Image size Tokens (standard level) Cost (Haiku 4.5, $1/1M)
200x200 px 64 ~$0.00006
1000x1000 px 1296 ~$0.0013
1920x1080 px 1560 (scaled down) ~$0.0016
Image size Tokens (high resolution) Cost (Opus 5.5, $4/1M)
200x200 px 64 ~$0.00026
1000x1000 px 1296 ~$0.0052
1920x1080 px 2691 (high resolution) ~$0.011

Claude 4.7 models and newer support high resolution. That's up to 3x more tokens per image, but they see fine detail better. If you don't need high precision, shrink the image before sending it.

🎨 Picture this: Eyesight that costs money. Imagine every time Claude looks at an image it costs money, like an X-ray at a clinic. A small image for reading text is a cheap X-ray (a fraction of a cent). A huge PNG for pixel-level analysis on a high-resolution model is a pricey MRI (around $0.02). Shrink the "scan" to the size you need.


Image quality tips (from Anthropic's documentation)

  • Format: JPEG, PNG, GIF or WebP. Animations aren't supported; only the first frame is used
  • Sharpness: the image should be sharp, not blurry or pixelated
  • Text: if the image has important text, make sure it's readable and not too small. Don't crop out visual context just to make the text bigger
  • Compression: lossy JPEG/WebP compression reduces request size and latency, but it can create artifacts. Check that text is still readable after compression
  • Resizing: keep in mind that the image may be scaled down automatically, which can make small text unreadable. It's better to resize it yourself ahead of time

Practical use cases

Analyzing UI screenshots:

Type this into the chat
Here's a screenshot of our checkout page [image].
Find UX problems: what might confuse the user, where there isn't 
enough contrast, which elements are placed in a confusing way.

Extracting data from documents:

Type this into the chat
Here's a photo of a restaurant receipt [image].
Extract: the restaurant name, the date, each item with its price, the total, 
and the tip if there is one. Return structured JSON.

Turning hand-drawn diagrams into code:

Type this into the chat
Here's a photo of a hand-drawn database diagram [image].
Convert it into SQL CREATE TABLE statements for PostgreSQL.

🎨 Picture this: a two-pass visual check is like a photographer who takes a shot and then looks at the camera screen. "The light's off, move a little to the right." Shoot → look → adjust → repeat. Without the second look, you print whatever came out on the first try.

Two-pass visual check (the Build → Screenshot → Review → Fix pattern):

A powerful pattern for UI work: first you write the code, then Claude checks the visual result:

python
# Step 1: Claude generates HTML/CSS
# Step 2: Take a screenshot of the page (with Playwright, Puppeteer, or Computer Use)
# Step 3: Send the screenshot to Claude and ask "what's wrong?"
# Step 4: Claude finds visual bugs and fixes the code

message = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=2048,
    messages=[{
        "role": "user",
        "content": [
            {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": screenshot_b64}},
            {"type": "text", "text": "This is a screenshot of a page I just generated. "
             "Find visual problems: misalignment, cut-off text, "
             "contrast issues, wrong spacing. Suggest CSS fixes."}
        ]
    }]
)

Automated design QA (comparing Figma with the build):

python
# Compare each component from Figma with a real screenshot
for component_name, figma_img, screenshot_img in components:
    result = client.messages.create(
        model="claude-opus-5-5",
        messages=[{
            "role": "user",
            "content": [
                {"type": "text", "text": "Image 1 (Figma mockup):"},
                {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": figma_img}},
                {"type": "text", "text": "Image 2 (build):"},
                {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": screenshot_img}},
                {"type": "text", "text": f"Component: {component_name}. Rate the match from 0 to 10 and describe the differences."}
            ]
        }]
    )

Practice

Assignment 1: Invoice analyzer

  1. Find or create a PNG image of an invoice. You can take a screenshot of any template online
  2. Write a Python script that reads the file, encodes it as base64 and sends it to Claude with a prompt (a prompt is your request to the AI): "Extract from this invoice: the document number, date, recipient, the list of line items (name, quantity, price) and the total amount. Return JSON."
  3. Run the script and check that the extracted data is correct
  4. Add error handling: when the file isn't found, when Claude couldn't extract the data
  5. Bonus: process several invoices in a loop and save the results to invoices.json

Assignment 2: A two-pass visual check in Claude Code

  1. Ask Claude Code to build an HTML page: "Create a landing page for a fitness app with a hero section, pricing and a CTA button"
  2. Open the result in your browser and take a screenshot (or use the Claude Code Read tool to view it)
  3. Show the screenshot to Claude: "Read /tmp/screenshot.png and find visual problems with this page"
  4. Claude will suggest fixes; apply them
  5. Repeat the loop: screenshot → review → fix, until you're happy with the result

Goal: understand how to pass images through the API and pull structured data out of visual sources. Get comfortable with the two-pass visual check pattern.


Tools and resources


Key takeaways

You can pass an image to the API in three ways: base64, URL, or the Files API (file_id). The type goes in the media_type field. Claude handles JPEG, PNG, GIF, WebP and PDF.

In Claude Code, images and PDFs open through the Read tool like regular files; no conversion needed. For large PDFs (more than 10 pages), give a page range.

Cost = ⌈width/28⌉ × ⌈height/28⌉ tokens. The maximum is 1568px on the long side (2576px for Claude 4.7 models and newer). Shrink images to the smallest size that works before sending.

Claude does NOT generate images; it only analyzes them. It doesn't identify people in photos. It can make mistakes on tiny, blurry or rotated images.

The two-pass visual check pattern (Build → Screenshot → Review → Fix) is a powerful tool for UI work, with Claude doing the visual verification.


Next lesson

→ AI Image Generation Pipeline: how to make images when Claude doesn't draw them. After that: What are Skills.

The mark stays in this browser only and is never sent anywhere. My progress