The gist
But with that power comes responsibility: give an assistant access to your computer and they may click the wrong button. In this lesson we won't just cover "how to turn on Computer Use." We'll cover how to build reliable, production-ready automation patterns that keep working even when something goes wrong.
Key concepts
- Computer Use API: Claude's ability to take screenshots and control the mouse and keyboard through special tools: the
computertoolset (in the current version,computer_toolset_20260801), plusbashandtext_editor - Screenshot-Analyze-Act loop: the core working loop: take a screenshot → understand what's on the screen → perform an action → check the result
- Verification-after-action pattern: after every click, take a confirming screenshot to make sure the action actually happened
- Retry loops with error recovery: smart retry loops that change strategy after an error instead of just repeating the same thing
- Environment isolation: Computer Use sees the user's entire desktop, so production needs a virtual machine or Docker with VNC
- Cost per operation: every screenshot-plus-analysis cycle uses tokens (the image plus the model's reply), so it costs noticeably more than a script; you need to know when it's worth it and when Playwright is the better choice
- Hybrid automation: combining Computer Use for the "hard" parts (login, non-standard elements, legacy UI) with Playwright for the "structured" parts (scraping data, clicking known selectors)
Theory
How the Computer Use API works
Computer Use isn't a separate model. It's a set of tools you pass to Claude when you call the API. Claude picks which tool to use depending on the task:
computer: take a screenshot, click, type text, press keys, scrollbash: run commands in the terminal (if allowed)text_editor: read and edit files
About tool versions. The examples below use the current computer_toolset_20260801 toolset: it doesn't need a beta header and doesn't take screen dimensions, and the model sends actions as separate calls (left_click, type, key, scroll, screenshot and others). Older models need the older tool version computer_20251124 with a beta header. For exact version names and the list of supported models, see the Computer Use documentation and the What's current page.
If you don't want to write code. Anthropic's ready-made products can also control a computer (for example, Claude Cowork and Claude Code on paid plans). What exactly is available on your plan is listed on the What's current page.
The basic loop looks like this:
Claude receives a task
→ Requests a screenshot
→ Analyzes what it sees
→ Decides which action to take
→ Performs the action (click, type, key)
→ Takes another screenshot and CHECKS the result
→ Repeats until the task is doneThe key detail: Claude itself decides when to take screenshots. Your job is to set the system up so it does this correctly and often enough.
Latency and realistic expectations
One "screenshot → analysis → action" cycle takes a few seconds, depending on screen size, the model and how complex the task is. For tasks with 20+ steps, that adds up to minutes of work.
Recommended settings to speed things up:
- Screen resolution: according to Anthropic's documentation, 1024×768 or 1280×720 works for general tasks, and 1280×800 or 1366×768 for web apps; it's best not to go above 1920×1080. There's no point sending a 4K monitor screenshot through the API: it's more expensive and slower
- Capture region: when possible, send only the part of the screen you need, not the whole desktop
- Headless VNC: in a Docker container with Xvfb you can guarantee a fixed resolution
Production pattern: a retry loop with smart recovery
The most common beginner mistake with Computer Use is having no retry logic. Interfaces change, elements load late, notifications pop up. The right pattern:
import anthropic
import base64
import time
from pathlib import Path
client = anthropic.Anthropic()
def take_screenshot() -> str:
"""
In production: takes a screenshot with scrot/PIL/mss and encodes it in base64.
This is a stub; in real use, replace it with your own implementation.
"""
# pip install mss Pillow
import mss
import io
from PIL import Image
with mss.mss() as sct:
monitor = {"top": 0, "left": 0, "width": 1280, "height": 800}
screenshot = sct.grab(monitor)
img = Image.frombytes("RGB", screenshot.size, screenshot.bgra, "raw", "BGRX")
buffer = io.BytesIO()
img.save(buffer, format="PNG")
return base64.standard_b64encode(buffer.getvalue()).decode("utf-8")
def image_block(b64: str) -> dict:
return {
"type": "image",
"source": {"type": "base64", "media_type": "image/png", "data": b64},
}
def computer_use_with_retry(
task: str,
max_steps: int = 30,
max_retries_per_step: int = 3,
pause_between_steps: float = 1.0,
) -> dict:
"""
Runs a Computer Use task with retry logic at every step.
Recovery pattern:
- If Claude reports an error → take a screenshot, pass along the context
- If an element isn't found → try scrolling or wait for it to load
- If errors keep piling up → escalate (stop, log)
"""
messages = []
# Current toolset: no beta header and no screen dimensions.
# The screenshots we return must fit within the model's size limits on their own.
tools = [{"type": "computer_toolset_20260801"}]
# Initial screenshot: "look at what's on the screen right now"
initial_screenshot = take_screenshot()
messages.append({
"role": "user",
"content": [
image_block(initial_screenshot),
{
"type": "text",
"text": f"""Here is the current state of the screen. Complete the following task:
{task}
IMPORTANT RULES:
1. After every click or text input, take a screenshot to check
2. If an element isn't visible, scroll first, then look for it
3. If something goes wrong, describe the problem in detail before the next attempt
4. When you finish the task, report TASK_COMPLETE and briefly describe what was done""",
},
],
})
step_count = 0
retry_context = []
while step_count < max_steps:
step_count += 1
try:
response = client.messages.create(
model="claude-opus-5-5", # current models: see the What's current page
max_tokens=4096,
tools=tools,
messages=messages,
)
# Check whether the task is finished
if response.stop_reason == "end_turn":
final_text = " ".join(
block.text for block in response.content
if block.type == "text"
)
if "TASK_COMPLETE" in final_text:
return {"status": "success", "steps": step_count, "summary": final_text}
# Finished without our marker: that's fine too
return {"status": "complete", "steps": step_count, "summary": final_text}
# Handle tool calls. The model may send several actions
# in a row: run them in order and stop at the first failure.
tool_results = []
failed = False
for block in response.content:
if block.type == "tool_use" and getattr(block, "toolset_name", None) == "computer":
result = {
"type": "tool_result",
"tool_use_id": block.id,
"toolset_name": "computer",
}
if failed:
result["is_error"] = True
result["content"] = "Not executed: an earlier computer action in this turn failed."
else:
try:
# block.name is the action itself: screenshot, left_click, type, key, scroll...
# In production, this is where your mouse/keyboard control code goes
execute_computer_action(block.name, block.input)
time.sleep(pause_between_steps)
if block.name in ("screenshot", "zoom"):
# KEY PATTERN: a fresh screenshot whenever the model asks for one.
# The "screenshot after every action" rule is set in the prompt above.
# For zoom, return the cropped region; here we return the full screen for brevity.
result["content"] = [image_block(take_screenshot())]
else:
result["content"] = [{"type": "text", "text": "OK"}]
except Exception as action_error:
failed = True
result["is_error"] = True
result["content"] = str(action_error)
tool_results.append(result)
# Add to the history and keep going
messages.append({"role": "assistant", "content": response.content})
if tool_results:
messages.append({"role": "user", "content": tool_results})
except Exception as e:
retry_context.append(str(e))
if len(retry_context) >= max_retries_per_step:
return {
"status": "error",
"steps": step_count,
"errors": retry_context,
}
# Add the error context and try again
error_screenshot = take_screenshot()
messages.append({
"role": "user",
"content": [
image_block(error_screenshot),
{"type": "text", "text": f"An error occurred: {str(e)}. Here is the current screen. Try a different approach."},
],
})
return {"status": "max_steps_reached", "steps": step_count}
# The model's key names (Return, Escape) differ from pyautogui's names (enter, esc)
KEY_MAP = {"return": "enter", "escape": "esc", "page_down": "pagedown", "page_up": "pageup"}
def execute_computer_action(name: str, params: dict) -> None:
"""
Stub: in real use, this is PyAutoGUI, xdotool or a native VNC client.
Action names and fields come from the computer tool documentation.
"""
# pip install pyautogui
import pyautogui
if name in ("screenshot", "zoom"):
return # we take the screenshot separately
elif name == "left_click":
x, y = params["coordinate"]
pyautogui.click(x, y)
elif name == "double_click":
x, y = params["coordinate"]
pyautogui.doubleClick(x, y)
elif name == "type":
pyautogui.write(params["text"], interval=0.05)
elif name == "key":
keys = [KEY_MAP.get(k.lower(), k.lower()) for k in params["text"].split("+")]
pyautogui.hotkey(*keys)
elif name == "scroll":
direction = params.get("scroll_direction", "down")
amount = params.get("scroll_amount", 3)
x, y = params.get("coordinate") or pyautogui.position()
pyautogui.scroll(amount if direction == "up" else -amount, x=x, y=y)
elif name == "wait":
time.sleep(params.get("duration", 1))
else:
raise ValueError(f"Action not supported: {name}")Multiple monitors and resolution normalization
Claude gets a screenshot and works with pixel coordinates. If the screen resolution changes, everything breaks. The fix:
# Always lock in a virtual resolution for Computer Use
VIRTUAL_WIDTH = 1280
VIRTUAL_HEIGHT = 800
# When capturing the real screen, scale down
# When passing Claude's coordinates back, scale up
def normalize_coordinates(x: int, y: int, real_width: int, real_height: int) -> tuple:
"""Convert Claude's virtual coordinates into real ones."""
real_x = int(x * real_width / VIRTUAL_WIDTH)
real_y = int(y * real_height / VIRTUAL_HEIGHT)
return real_x, real_yMultiple monitors are a separate story. The simplest approach: run the task on one specific monitor using an offset (monitor = {"top": 0, "left": 1920, ...} for the second monitor).
Native desktop apps: where Computer Use is irreplaceable
Playwright, Selenium and API integrations all work with web interfaces. But there's a huge class of tasks with no web involved:
- Outdated accounting software (old desktop bookkeeping programs, old ERPs with no REST API)
- Xcode: building an iOS project, automating UI tests
- Figma desktop: batch operations on components
- Specialized B2B software: customs declarations, desktop banking clients
- Desktop games: automating repetitive actions
For all of these, Computer Use is the only automation tool that doesn't require writing custom native hooks.
Real-world case: automating bookkeeping in legacy software
The task: every day, download a report from a Windows-only tax reporting program (no API), parse the data and upload it to Google Sheets.
The Computer Use solution:
- Start a Windows VM with VNC
- Open the program through Computer Use
- Claude navigates the menus, picks the period and exports the file
- Pull the file out of the VM with bash, parse it and send it to Sheets
This can't be done with Playwright or regular automation: the program has no web interface.
Headless vs. headed: what to give up
| Mode | Pros | Cons | When to use |
|---|---|---|---|
| Headed (real screen) | You can debug visually | Needs a monitor or an X server | Development, testing |
| Headless with Xvfb | Runs in Docker without a monitor | Can't debug without VNC | Production server |
| VNC in Docker | You can watch remotely | Extra complexity | CI/CD + debugging |
# Run headless with the option to watch over VNC
docker run -d \
-e DISPLAY=:1 \
-p 5900:5900 \
--name cu-sandbox \
my-computer-use-image
# Inside the container
Xvfb :1 -screen 0 1280x800x24 &
x11vnc -display :1 -nopw -listen 0.0.0.0 -forever &Cost: when Computer Use is worth it
| Task | Best tool | Why |
|---|---|---|
| Clicking known HTML selectors | Playwright | Fast, nearly free (no model tokens), reliable |
| Non-standard interface without stable selectors (canvas, custom widgets) | Computer Use | Playwright can't see visually |
| Login via SSO / OAuth with 2FA | Computer Use | No API for this flow |
| Scraping data from a table | Playwright + CSS | Structured data |
| Working in a native desktop app | Computer Use | No alternative |
| Testing UI across browsers | Playwright | Built-in cross-browser support |
How to estimate the cost of one Computer Use operation: number of steps × (screenshot tokens + model reply tokens) × the model's price per million tokens. The context grows with every step, so long tasks get more expensive faster than you'd expect. Model prices: What's current. Before running a task on a schedule, run it a few times and look at the token usage (the usage field in the API response).
For tasks with no alternative, paying for the tokens is justified. For tasks that have an API or can use Playwright, it's a noticeable overpayment.
Security: what Computer Use can see
Computer Use has access to everything the user session can see:
- All open browser tabs
- Files on the desktop
- The clipboard
- All running applications
That means: never run Computer Use in the same session where your password manager, personal browser or company systems are open. Use an isolated VM or a Docker container with a clean user account.
Anthropic's documentation also recommends: don't give the model access to sensitive data (logins, passwords), limit internet access to an allowlist of domains, and have a human confirm actions with real-world consequences (payments, accepting terms, accepting cookies). A page on the screen may contain hidden instructions for the model (prompt injection): more on this in the lesson Defending against prompt injection.
# The right production architecture
# 1. A separate Docker container with VNC
# 2. Inside the container, a clean user with no access to production data
# 3. Only the apps you need are installed
# 4. Results are passed through a volume mount, not through the clipboardHybrid pattern: Computer Use + Playwright
The most powerful approach in real projects is to use each tool for what it does best:
from playwright.async_api import async_playwright
async def hybrid_automation():
# Step 1: Computer Use for the tricky login (SSO + 2FA)
login_result = computer_use_with_retry(
task="Open example.com, click 'Sign in with company account', "
"enter the login [email protected], wait for the 2FA text message and enter the code",
max_steps=20
)
# Step 2: Playwright picks up the already authenticated session
# (we pass cookies or storage state from the browser)
async with async_playwright() as p:
browser = await p.chromium.connect_over_cdp("http://localhost:9222")
page = browser.contexts[0].pages[0] # take the page that's already open
# Now fast, structured work through selectors
rows = await page.query_selector_all("table.reports tr")
data = []
for row in rows:
cells = await row.query_selector_all("td")
data.append([await cell.inner_text() for cell in cells])
return dataError handling: what to do when Claude clicks the wrong thing
Three levels of problems and their fixes:
The click missed the element → the verification screenshot shows that nothing changed → Claude tries again with corrected coordinates
The element you need never appeared → add waiting logic to the retry loop: "if the element isn't visible, scroll the page or wait 2 seconds"
An unexpected pop-up → Claude needs to be able to handle dialogs, notifications and cookie banners. Add this to the system prompt: "If any modal window appears, close it before continuing with the main task"
Practice
Task: Automate a daily report export from a desktop app.
Step 1: Set up an isolated environment
# Install dependencies
pip install anthropic mss Pillow pyautogui
# For Linux/Docker: install Xvfb + x11vnc
# sudo apt-get install xvfb x11vnc
# Start a virtual display (only for Linux without a monitor)
export DISPLAY=:1
Xvfb :1 -screen 0 1280x800x24 &Step 2: Create the file cu_screenshot.py
import mss
import io
import base64
from PIL import Image
VIRTUAL_WIDTH = 1280
VIRTUAL_HEIGHT = 800
def capture_screen(region=None) -> str:
"""Takes a screenshot and returns it as base64."""
with mss.mss() as sct:
monitor = region or {"top": 0, "left": 0, "width": VIRTUAL_WIDTH, "height": VIRTUAL_HEIGHT}
screenshot = sct.grab(monitor)
img = Image.frombytes("RGB", screenshot.size, screenshot.bgra, "raw", "BGRX")
img = img.resize((VIRTUAL_WIDTH, VIRTUAL_HEIGHT))
buffer = io.BytesIO()
img.save(buffer, format="PNG", optimize=True)
return base64.standard_b64encode(buffer.getvalue()).decode("utf-8")Step 3: Build the action executor
import pyautogui
import time
pyautogui.FAILSAFE = True # Mouse to a corner = stop
# The model's key names (Return, Escape) differ from pyautogui's names (enter, esc)
KEY_MAP = {"return": "enter", "escape": "esc"}
def execute_action(name: str, params: dict) -> None:
"""Performs one model action. Raises an exception on failure: the calling code returns is_error."""
if name in ("screenshot", "zoom"):
return # the calling code takes the screenshot
if name == "left_click":
x, y = params["coordinate"]
pyautogui.click(x, y)
elif name == "double_click":
x, y = params["coordinate"]
pyautogui.doubleClick(x, y)
elif name == "type":
time.sleep(0.2) # Short pause before typing
pyautogui.write(params["text"], interval=0.03)
elif name == "key":
keys = [KEY_MAP.get(k.lower(), k.lower()) for k in params["text"].split("+")]
pyautogui.hotkey(*keys)
elif name == "scroll":
direction = params.get("scroll_direction", "down")
amount = params.get("scroll_amount", 3)
x, y = params.get("coordinate") or pyautogui.position()
pyautogui.scroll(-amount if direction == "down" else amount, x=x, y=y)
elif name == "wait":
time.sleep(params.get("duration", 1))
else:
raise ValueError(f"Action not supported: {name}")
time.sleep(0.5) # Wait for the UI to respondStep 4: Run the task with verification
import anthropic
import time
from cu_screenshot import capture_screen
from executor import execute_action
def screenshot_block() -> dict:
return {
"type": "image",
"source": {"type": "base64", "media_type": "image/png", "data": capture_screen()},
}
def run_desktop_task(task_description: str, app_name: str):
client = anthropic.Anthropic()
# No beta header and no screen dimensions (current toolset)
tools = [{"type": "computer_toolset_20260801"}]
system_prompt = f"""You are automating a task in the {app_name} application.
RULES:
- After every action, take a screenshot to check
- If you see a modal window or a notification, close it
- If an element isn't found, scroll the page, wait 2 seconds and try again
- When the task is done, write TASK_COMPLETE and describe what you did
- If the task is impossible, write TASK_FAILED and explain why"""
messages = [{
"role": "user",
"content": [
screenshot_block(),
{"type": "text", "text": f"Complete the task: {task_description}"},
],
}]
for step in range(40): # 40 steps max
response = client.messages.create(
model="claude-opus-5-5", # current models: see the What's current page
max_tokens=4096,
system=system_prompt,
tools=tools,
messages=messages,
)
if response.stop_reason == "end_turn":
final = " ".join(b.text for b in response.content if b.type == "text")
print(f"Finished in {step+1} steps: {final}")
return "TASK_COMPLETE" in final
tool_results = []
failed = False
for block in response.content:
if block.type == "tool_use" and getattr(block, "toolset_name", None) == "computer":
result = {"type": "tool_result", "tool_use_id": block.id, "toolset_name": "computer"}
if failed:
result["is_error"] = True
result["content"] = "Not executed: an earlier computer action in this turn failed."
else:
try:
execute_action(block.name, block.input)
time.sleep(1.0)
if block.name in ("screenshot", "zoom"):
result["content"] = [screenshot_block()]
else:
result["content"] = [{"type": "text", "text": "OK"}]
except Exception as e:
failed = True
print(f"Error performing action {block.name}: {e}")
result["is_error"] = True
result["content"] = str(e)
tool_results.append(result)
messages.append({"role": "assistant", "content": response.content})
if tool_results:
messages.append({"role": "user", "content": tool_results})
print("Maximum number of steps exceeded")
return False
# Usage
run_desktop_task(
task_description="Open the File → Reports → Daily menu. Select yesterday's date. Click Export → CSV. Save to the /tmp/reports/ folder",
app_name="LegacyAccountingApp"
)Step 5: Add logging and monitoring
import json
from datetime import datetime
from pathlib import Path
def log_session(task: str, success: bool, steps: int, errors: list):
log_entry = {
"timestamp": datetime.now().isoformat(),
"task": task[:100],
"success": success,
"steps": steps,
"errors": errors,
}
log_path = Path("logs/computer_use.jsonl")
log_path.parent.mkdir(exist_ok=True)
with open(log_path, "a") as f:
f.write(json.dumps(log_entry, ensure_ascii=False) + "\n")Tools and resources
- Anthropic Computer Use API: official documentation covering tool versions, actions and security
- anthropic/computer-use-demo: a ready-made Docker image from Anthropic for a quick start (the example may lag behind the current API version, so check it against the documentation)
- mss: fast screenshots in Python (faster than PIL)
- pyautogui: mouse and keyboard control in Python (cross-platform)
- Playwright: for hybrid scenarios (CU for login, Playwright for data)
- xdotool: a Linux alternative to pyautogui, more reliable for headless setups
- Xvfb: a virtual X server for headless Linux
- VNC + noVNC: view a virtual display remotely in your browser
Key takeaways
"Computer Use doesn't replace Playwright. It fills in for tasks where there's no other way: native apps, complex SSO flows, legacy UI with no API"
"A verification screenshot after every action isn't optional. It's a required part of reliable automation. Without it, Claude is working blind"
"Estimate the cost up front: every step is a screenshot plus a model reply, paid in tokens. If there's an API or Playwright, use it. Computer Use is justified only where there's no alternative"
Next lesson
→ Voice AI Agents: Vapi, Bland.ai and phone agents
The mark stays in this browser only and is never sent anywhere. My progress