Library · Autonomous and multi-agent systems

Advanced Computer Use: patterns for complex automation

Engineer65 minUpdated: October 2026
75 of 105 in the library

Module: 18. Advanced Orchestration | Time: about 25 min theory + 40 min practice


The gist

🎨 Picture this: Until now you had a very smart assistant who could only read documents and tell you what to do. Now you've given that assistant eyes and hands and sat them down at your computer. They see the screen the same way you do. They can click buttons, type text and drag files around. That's Computer Use: Claude goes from being a "brain" to being a full operator at a workstation.

But with that power comes responsibility: give an assistant access to your computer and they may click the wrong button. In this lesson we won't just cover "how to turn on Computer Use." We'll cover how to build reliable, production-ready automation patterns that keep working even when something goes wrong.


Key concepts

  • Computer Use API: Claude's ability to take screenshots and control the mouse and keyboard through special tools: the computer toolset (in the current version, computer_toolset_20260801), plus bash and text_editor
  • Screenshot-Analyze-Act loop: the core working loop: take a screenshot → understand what's on the screen → perform an action → check the result
  • Verification-after-action pattern: after every click, take a confirming screenshot to make sure the action actually happened
  • Retry loops with error recovery: smart retry loops that change strategy after an error instead of just repeating the same thing
  • Environment isolation: Computer Use sees the user's entire desktop, so production needs a virtual machine or Docker with VNC
  • Cost per operation: every screenshot-plus-analysis cycle uses tokens (the image plus the model's reply), so it costs noticeably more than a script; you need to know when it's worth it and when Playwright is the better choice
  • Hybrid automation: combining Computer Use for the "hard" parts (login, non-standard elements, legacy UI) with Playwright for the "structured" parts (scraping data, clicking known selectors)

Theory

How the Computer Use API works

Computer Use isn't a separate model. It's a set of tools you pass to Claude when you call the API. Claude picks which tool to use depending on the task:

  • computer: take a screenshot, click, type text, press keys, scroll
  • bash: run commands in the terminal (if allowed)
  • text_editor: read and edit files

About tool versions. The examples below use the current computer_toolset_20260801 toolset: it doesn't need a beta header and doesn't take screen dimensions, and the model sends actions as separate calls (left_click, type, key, scroll, screenshot and others). Older models need the older tool version computer_20251124 with a beta header. For exact version names and the list of supported models, see the Computer Use documentation and the What's current page.

If you don't want to write code. Anthropic's ready-made products can also control a computer (for example, Claude Cowork and Claude Code on paid plans). What exactly is available on your plan is listed on the What's current page.

The basic loop looks like this:

Code
Claude receives a task
→ Requests a screenshot
→ Analyzes what it sees
→ Decides which action to take
→ Performs the action (click, type, key)
→ Takes another screenshot and CHECKS the result
→ Repeats until the task is done

The key detail: Claude itself decides when to take screenshots. Your job is to set the system up so it does this correctly and often enough.

Latency and realistic expectations

One "screenshot → analysis → action" cycle takes a few seconds, depending on screen size, the model and how complex the task is. For tasks with 20+ steps, that adds up to minutes of work.

Recommended settings to speed things up:

  • Screen resolution: according to Anthropic's documentation, 1024×768 or 1280×720 works for general tasks, and 1280×800 or 1366×768 for web apps; it's best not to go above 1920×1080. There's no point sending a 4K monitor screenshot through the API: it's more expensive and slower
  • Capture region: when possible, send only the part of the screen you need, not the whole desktop
  • Headless VNC: in a Docker container with Xvfb you can guarantee a fixed resolution

Production pattern: a retry loop with smart recovery

The most common beginner mistake with Computer Use is having no retry logic. Interfaces change, elements load late, notifications pop up. The right pattern:

python
import anthropic
import base64
import time
from pathlib import Path

client = anthropic.Anthropic()

def take_screenshot() -> str:
    """
    In production: takes a screenshot with scrot/PIL/mss and encodes it in base64.
    This is a stub; in real use, replace it with your own implementation.
    """
    # pip install mss Pillow
    import mss
    import io
    from PIL import Image
    
    with mss.mss() as sct:
        monitor = {"top": 0, "left": 0, "width": 1280, "height": 800}
        screenshot = sct.grab(monitor)
        img = Image.frombytes("RGB", screenshot.size, screenshot.bgra, "raw", "BGRX")
        buffer = io.BytesIO()
        img.save(buffer, format="PNG")
        return base64.standard_b64encode(buffer.getvalue()).decode("utf-8")


def image_block(b64: str) -> dict:
    return {
        "type": "image",
        "source": {"type": "base64", "media_type": "image/png", "data": b64},
    }


def computer_use_with_retry(
    task: str,
    max_steps: int = 30,
    max_retries_per_step: int = 3,
    pause_between_steps: float = 1.0,
) -> dict:
    """
    Runs a Computer Use task with retry logic at every step.
    
    Recovery pattern:
    - If Claude reports an error → take a screenshot, pass along the context
    - If an element isn't found → try scrolling or wait for it to load
    - If errors keep piling up → escalate (stop, log)
    """
    
    messages = []
    # Current toolset: no beta header and no screen dimensions.
    # The screenshots we return must fit within the model's size limits on their own.
    tools = [{"type": "computer_toolset_20260801"}]
    
    # Initial screenshot: "look at what's on the screen right now"
    initial_screenshot = take_screenshot()
    
    messages.append({
        "role": "user",
        "content": [
            image_block(initial_screenshot),
            {
                "type": "text",
                "text": f"""Here is the current state of the screen. Complete the following task:

{task}

IMPORTANT RULES:
1. After every click or text input, take a screenshot to check
2. If an element isn't visible, scroll first, then look for it
3. If something goes wrong, describe the problem in detail before the next attempt
4. When you finish the task, report TASK_COMPLETE and briefly describe what was done""",
            },
        ],
    })
    
    step_count = 0
    retry_context = []
    
    while step_count < max_steps:
        step_count += 1
        
        try:
            response = client.messages.create(
                model="claude-opus-5-5",   # current models: see the What's current page
                max_tokens=4096,
                tools=tools,
                messages=messages,
            )
            
            # Check whether the task is finished
            if response.stop_reason == "end_turn":
                final_text = " ".join(
                    block.text for block in response.content 
                    if block.type == "text"
                )
                if "TASK_COMPLETE" in final_text:
                    return {"status": "success", "steps": step_count, "summary": final_text}
                # Finished without our marker: that's fine too
                return {"status": "complete", "steps": step_count, "summary": final_text}
            
            # Handle tool calls. The model may send several actions
            # in a row: run them in order and stop at the first failure.
            tool_results = []
            failed = False
            for block in response.content:
                if block.type == "tool_use" and getattr(block, "toolset_name", None) == "computer":
                    result = {
                        "type": "tool_result",
                        "tool_use_id": block.id,
                        "toolset_name": "computer",
                    }
                    if failed:
                        result["is_error"] = True
                        result["content"] = "Not executed: an earlier computer action in this turn failed."
                    else:
                        try:
                            # block.name is the action itself: screenshot, left_click, type, key, scroll...
                            # In production, this is where your mouse/keyboard control code goes
                            execute_computer_action(block.name, block.input)
                            time.sleep(pause_between_steps)
                            if block.name in ("screenshot", "zoom"):
                                # KEY PATTERN: a fresh screenshot whenever the model asks for one.
                                # The "screenshot after every action" rule is set in the prompt above.
                                # For zoom, return the cropped region; here we return the full screen for brevity.
                                result["content"] = [image_block(take_screenshot())]
                            else:
                                result["content"] = [{"type": "text", "text": "OK"}]
                        except Exception as action_error:
                            failed = True
                            result["is_error"] = True
                            result["content"] = str(action_error)
                    tool_results.append(result)
            
            # Add to the history and keep going
            messages.append({"role": "assistant", "content": response.content})
            if tool_results:
                messages.append({"role": "user", "content": tool_results})
                
        except Exception as e:
            retry_context.append(str(e))
            if len(retry_context) >= max_retries_per_step:
                return {
                    "status": "error",
                    "steps": step_count,
                    "errors": retry_context,
                }
            # Add the error context and try again
            error_screenshot = take_screenshot()
            messages.append({
                "role": "user",
                "content": [
                    image_block(error_screenshot),
                    {"type": "text", "text": f"An error occurred: {str(e)}. Here is the current screen. Try a different approach."},
                ],
            })
    
    return {"status": "max_steps_reached", "steps": step_count}


# The model's key names (Return, Escape) differ from pyautogui's names (enter, esc)
KEY_MAP = {"return": "enter", "escape": "esc", "page_down": "pagedown", "page_up": "pageup"}


def execute_computer_action(name: str, params: dict) -> None:
    """
    Stub: in real use, this is PyAutoGUI, xdotool or a native VNC client.
    Action names and fields come from the computer tool documentation.
    """
    # pip install pyautogui
    import pyautogui
    
    if name in ("screenshot", "zoom"):
        return  # we take the screenshot separately
    elif name == "left_click":
        x, y = params["coordinate"]
        pyautogui.click(x, y)
    elif name == "double_click":
        x, y = params["coordinate"]
        pyautogui.doubleClick(x, y)
    elif name == "type":
        pyautogui.write(params["text"], interval=0.05)
    elif name == "key":
        keys = [KEY_MAP.get(k.lower(), k.lower()) for k in params["text"].split("+")]
        pyautogui.hotkey(*keys)
    elif name == "scroll":
        direction = params.get("scroll_direction", "down")
        amount = params.get("scroll_amount", 3)
        x, y = params.get("coordinate") or pyautogui.position()
        pyautogui.scroll(amount if direction == "up" else -amount, x=x, y=y)
    elif name == "wait":
        time.sleep(params.get("duration", 1))
    else:
        raise ValueError(f"Action not supported: {name}")

Multiple monitors and resolution normalization

🎨 Picture this: You're telling a friend how to find a button on your screen. You say, "click in the top right corner." But if your friend's screen is twice as big, your coordinates will send them somewhere else entirely.

Claude gets a screenshot and works with pixel coordinates. If the screen resolution changes, everything breaks. The fix:

python
# Always lock in a virtual resolution for Computer Use
VIRTUAL_WIDTH = 1280
VIRTUAL_HEIGHT = 800

# When capturing the real screen, scale down
# When passing Claude's coordinates back, scale up

def normalize_coordinates(x: int, y: int, real_width: int, real_height: int) -> tuple:
    """Convert Claude's virtual coordinates into real ones."""
    real_x = int(x * real_width / VIRTUAL_WIDTH)
    real_y = int(y * real_height / VIRTUAL_HEIGHT)
    return real_x, real_y

Multiple monitors are a separate story. The simplest approach: run the task on one specific monitor using an offset (monitor = {"top": 0, "left": 1920, ...} for the second monitor).

Native desktop apps: where Computer Use is irreplaceable

Playwright, Selenium and API integrations all work with web interfaces. But there's a huge class of tasks with no web involved:

  • Outdated accounting software (old desktop bookkeeping programs, old ERPs with no REST API)
  • Xcode: building an iOS project, automating UI tests
  • Figma desktop: batch operations on components
  • Specialized B2B software: customs declarations, desktop banking clients
  • Desktop games: automating repetitive actions

For all of these, Computer Use is the only automation tool that doesn't require writing custom native hooks.

Real-world case: automating bookkeeping in legacy software

The task: every day, download a report from a Windows-only tax reporting program (no API), parse the data and upload it to Google Sheets.

The Computer Use solution:

  1. Start a Windows VM with VNC
  2. Open the program through Computer Use
  3. Claude navigates the menus, picks the period and exports the file
  4. Pull the file out of the VM with bash, parse it and send it to Sheets

This can't be done with Playwright or regular automation: the program has no web interface.

Headless vs. headed: what to give up

Mode Pros Cons When to use
Headed (real screen) You can debug visually Needs a monitor or an X server Development, testing
Headless with Xvfb Runs in Docker without a monitor Can't debug without VNC Production server
VNC in Docker You can watch remotely Extra complexity CI/CD + debugging
bash
# Run headless with the option to watch over VNC
docker run -d \
  -e DISPLAY=:1 \
  -p 5900:5900 \
  --name cu-sandbox \
  my-computer-use-image

# Inside the container
Xvfb :1 -screen 0 1280x800x24 &
x11vnc -display :1 -nopw -listen 0.0.0.0 -forever &

Cost: when Computer Use is worth it

🎨 Picture this: Computer Use is like hiring a highly skilled specialist by the hour for a task an intern could automate. Sometimes you do need the specialist, but it's important to know when.

Task Best tool Why
Clicking known HTML selectors Playwright Fast, nearly free (no model tokens), reliable
Non-standard interface without stable selectors (canvas, custom widgets) Computer Use Playwright can't see visually
Login via SSO / OAuth with 2FA Computer Use No API for this flow
Scraping data from a table Playwright + CSS Structured data
Working in a native desktop app Computer Use No alternative
Testing UI across browsers Playwright Built-in cross-browser support

How to estimate the cost of one Computer Use operation: number of steps × (screenshot tokens + model reply tokens) × the model's price per million tokens. The context grows with every step, so long tasks get more expensive faster than you'd expect. Model prices: What's current. Before running a task on a schedule, run it a few times and look at the token usage (the usage field in the API response).

For tasks with no alternative, paying for the tokens is justified. For tasks that have an API or can use Playwright, it's a noticeable overpayment.

Security: what Computer Use can see

Computer Use has access to everything the user session can see:

  • All open browser tabs
  • Files on the desktop
  • The clipboard
  • All running applications

That means: never run Computer Use in the same session where your password manager, personal browser or company systems are open. Use an isolated VM or a Docker container with a clean user account.

Anthropic's documentation also recommends: don't give the model access to sensitive data (logins, passwords), limit internet access to an allowlist of domains, and have a human confirm actions with real-world consequences (payments, accepting terms, accepting cookies). A page on the screen may contain hidden instructions for the model (prompt injection): more on this in the lesson Defending against prompt injection.

python
# The right production architecture
# 1. A separate Docker container with VNC
# 2. Inside the container, a clean user with no access to production data
# 3. Only the apps you need are installed
# 4. Results are passed through a volume mount, not through the clipboard

Hybrid pattern: Computer Use + Playwright

The most powerful approach in real projects is to use each tool for what it does best:

python
from playwright.async_api import async_playwright

async def hybrid_automation():
    # Step 1: Computer Use for the tricky login (SSO + 2FA)
    login_result = computer_use_with_retry(
        task="Open example.com, click 'Sign in with company account', "
             "enter the login [email protected], wait for the 2FA text message and enter the code",
        max_steps=20
    )
    
    # Step 2: Playwright picks up the already authenticated session
    # (we pass cookies or storage state from the browser)
    async with async_playwright() as p:
        browser = await p.chromium.connect_over_cdp("http://localhost:9222")
        page = browser.contexts[0].pages[0]  # take the page that's already open
        
        # Now fast, structured work through selectors
        rows = await page.query_selector_all("table.reports tr")
        data = []
        for row in rows:
            cells = await row.query_selector_all("td")
            data.append([await cell.inner_text() for cell in cells])
        
        return data

Error handling: what to do when Claude clicks the wrong thing

Three levels of problems and their fixes:

  1. The click missed the element → the verification screenshot shows that nothing changed → Claude tries again with corrected coordinates

  2. The element you need never appeared → add waiting logic to the retry loop: "if the element isn't visible, scroll the page or wait 2 seconds"

  3. An unexpected pop-up → Claude needs to be able to handle dialogs, notifications and cookie banners. Add this to the system prompt: "If any modal window appears, close it before continuing with the main task"


Practice

Task: Automate a daily report export from a desktop app.

Step 1: Set up an isolated environment

bash
# Install dependencies
pip install anthropic mss Pillow pyautogui

# For Linux/Docker: install Xvfb + x11vnc
# sudo apt-get install xvfb x11vnc

# Start a virtual display (only for Linux without a monitor)
export DISPLAY=:1
Xvfb :1 -screen 0 1280x800x24 &

Step 2: Create the file cu_screenshot.py

python
import mss
import io
import base64
from PIL import Image

VIRTUAL_WIDTH = 1280
VIRTUAL_HEIGHT = 800

def capture_screen(region=None) -> str:
    """Takes a screenshot and returns it as base64."""
    with mss.mss() as sct:
        monitor = region or {"top": 0, "left": 0, "width": VIRTUAL_WIDTH, "height": VIRTUAL_HEIGHT}
        screenshot = sct.grab(monitor)
        img = Image.frombytes("RGB", screenshot.size, screenshot.bgra, "raw", "BGRX")
        img = img.resize((VIRTUAL_WIDTH, VIRTUAL_HEIGHT))
        buffer = io.BytesIO()
        img.save(buffer, format="PNG", optimize=True)
        return base64.standard_b64encode(buffer.getvalue()).decode("utf-8")

Step 3: Build the action executor

python
import pyautogui
import time

pyautogui.FAILSAFE = True  # Mouse to a corner = stop

# The model's key names (Return, Escape) differ from pyautogui's names (enter, esc)
KEY_MAP = {"return": "enter", "escape": "esc"}

def execute_action(name: str, params: dict) -> None:
    """Performs one model action. Raises an exception on failure: the calling code returns is_error."""
    if name in ("screenshot", "zoom"):
        return  # the calling code takes the screenshot
    if name == "left_click":
        x, y = params["coordinate"]
        pyautogui.click(x, y)
    elif name == "double_click":
        x, y = params["coordinate"]
        pyautogui.doubleClick(x, y)
    elif name == "type":
        time.sleep(0.2)  # Short pause before typing
        pyautogui.write(params["text"], interval=0.03)
    elif name == "key":
        keys = [KEY_MAP.get(k.lower(), k.lower()) for k in params["text"].split("+")]
        pyautogui.hotkey(*keys)
    elif name == "scroll":
        direction = params.get("scroll_direction", "down")
        amount = params.get("scroll_amount", 3)
        x, y = params.get("coordinate") or pyautogui.position()
        pyautogui.scroll(-amount if direction == "down" else amount, x=x, y=y)
    elif name == "wait":
        time.sleep(params.get("duration", 1))
    else:
        raise ValueError(f"Action not supported: {name}")
    time.sleep(0.5)  # Wait for the UI to respond

Step 4: Run the task with verification

python
import anthropic
import time
from cu_screenshot import capture_screen
from executor import execute_action

def screenshot_block() -> dict:
    return {
        "type": "image",
        "source": {"type": "base64", "media_type": "image/png", "data": capture_screen()},
    }

def run_desktop_task(task_description: str, app_name: str):
    client = anthropic.Anthropic()
    
    # No beta header and no screen dimensions (current toolset)
    tools = [{"type": "computer_toolset_20260801"}]
    
    system_prompt = f"""You are automating a task in the {app_name} application.
    
RULES:
- After every action, take a screenshot to check
- If you see a modal window or a notification, close it
- If an element isn't found, scroll the page, wait 2 seconds and try again
- When the task is done, write TASK_COMPLETE and describe what you did
- If the task is impossible, write TASK_FAILED and explain why"""
    
    messages = [{
        "role": "user",
        "content": [
            screenshot_block(),
            {"type": "text", "text": f"Complete the task: {task_description}"},
        ],
    }]
    
    for step in range(40):  # 40 steps max
        response = client.messages.create(
            model="claude-opus-5-5",   # current models: see the What's current page
            max_tokens=4096,
            system=system_prompt,
            tools=tools,
            messages=messages,
        )
        
        if response.stop_reason == "end_turn":
            final = " ".join(b.text for b in response.content if b.type == "text")
            print(f"Finished in {step+1} steps: {final}")
            return "TASK_COMPLETE" in final
        
        tool_results = []
        failed = False
        for block in response.content:
            if block.type == "tool_use" and getattr(block, "toolset_name", None) == "computer":
                result = {"type": "tool_result", "tool_use_id": block.id, "toolset_name": "computer"}
                if failed:
                    result["is_error"] = True
                    result["content"] = "Not executed: an earlier computer action in this turn failed."
                else:
                    try:
                        execute_action(block.name, block.input)
                        time.sleep(1.0)
                        if block.name in ("screenshot", "zoom"):
                            result["content"] = [screenshot_block()]
                        else:
                            result["content"] = [{"type": "text", "text": "OK"}]
                    except Exception as e:
                        failed = True
                        print(f"Error performing action {block.name}: {e}")
                        result["is_error"] = True
                        result["content"] = str(e)
                tool_results.append(result)
        
        messages.append({"role": "assistant", "content": response.content})
        if tool_results:
            messages.append({"role": "user", "content": tool_results})
    
    print("Maximum number of steps exceeded")
    return False

# Usage
run_desktop_task(
    task_description="Open the File → Reports → Daily menu. Select yesterday's date. Click Export → CSV. Save to the /tmp/reports/ folder",
    app_name="LegacyAccountingApp"
)

Step 5: Add logging and monitoring

python
import json
from datetime import datetime
from pathlib import Path

def log_session(task: str, success: bool, steps: int, errors: list):
    log_entry = {
        "timestamp": datetime.now().isoformat(),
        "task": task[:100],
        "success": success,
        "steps": steps,
        "errors": errors,
    }
    log_path = Path("logs/computer_use.jsonl")
    log_path.parent.mkdir(exist_ok=True)
    with open(log_path, "a") as f:
        f.write(json.dumps(log_entry, ensure_ascii=False) + "\n")

Tools and resources

  • Anthropic Computer Use API: official documentation covering tool versions, actions and security
  • anthropic/computer-use-demo: a ready-made Docker image from Anthropic for a quick start (the example may lag behind the current API version, so check it against the documentation)
  • mss: fast screenshots in Python (faster than PIL)
  • pyautogui: mouse and keyboard control in Python (cross-platform)
  • Playwright: for hybrid scenarios (CU for login, Playwright for data)
  • xdotool: a Linux alternative to pyautogui, more reliable for headless setups
  • Xvfb: a virtual X server for headless Linux
  • VNC + noVNC: view a virtual display remotely in your browser

Key takeaways

"Computer Use doesn't replace Playwright. It fills in for tasks where there's no other way: native apps, complex SSO flows, legacy UI with no API"

"A verification screenshot after every action isn't optional. It's a required part of reliable automation. Without it, Claude is working blind"

"Estimate the cost up front: every step is a screenshot plus a model reply, paid in tokens. If there's an API or Playwright, use it. Computer Use is justified only where there's no alternative"


Next lesson

→ Voice AI Agents: Vapi, Bland.ai and phone agents

The mark stays in this browser only and is never sent anywhere. My progress