Library · Autonomous and multi-agent systems

Browser agents: Browserbase, Stagehand and web automation

Engineer60 minUpdated: October 2026
73 of 105 in the library

Module: 18. Advanced Orchestration | Time: ~25 min theory + 35 min practice


The gist

🎨 Picture this: you have an invisible employee who never sleeps, never gets tired and never asks for a paycheck. You tell them: "Go to the competitor's website, find all the laptop prices, save them to a spreadsheet and send them to me at 9 a.m." And they do it, every day, around the clock, without mistakes or complaints.

That's a browser agent: an AI system that drives a real browser, reads pages the way a person does and decides on its next action based on what it sees. Not just a script with fixed CSS selectors, but real intelligence that understands what a page means.

In this lesson we'll look at how to build agents like this with Browserbase and Stagehand, see how they differ from classic automation, and write a working agent that monitors competitor prices.


Key concepts

  • Browser agent vs. classic automation: AI understands what a page means instead of just looking for CSS selectors
  • Browserbase: cloud browser infrastructure for AI agents, with anti-detection features
  • Stagehand: an SDK from Browserbase built on top of Playwright, with LLM-driven methods act(), extract() and observe()
  • Session replay: recording and replaying sessions to debug agents
  • Anti-detection techniques: stealth mode, residential proxies, getting past CAPTCHAs
  • Scraping ethics: public vs. private data, Terms of Service, PII
  • Monetization: data collection and monitoring services (a service format, not a guaranteed income)

Theory

Why browser agents aren't Selenium 2.0

When Selenium appeared in 2004, web automation was built on one principle: a programmer manually finds the CSS selector or XPath to the element they need and tells the script, "click #buy-button." That worked until a site changed its layout. As soon as the developers renamed the button's class, the script broke.

Puppeteer (2017) and Playwright (2020) made this approach faster and more reliable, but the principle stayed the same: rigid instructions based on the DOM structure. If the page changes, the automation breaks.

A browser agent thinks differently. Instead of "click the element with id=submit," it gets the task "place the order" and figures out on its own what that takes. It reads the text on the page, understands the context, sees that the button says "Place order" or "Checkout," and clicks it. If the layout changes, the agent simply reworks its plan.

🎨 Picture this: a classic script is a blindfolded worker given exact coordinates: "turn 45 degrees, take three steps forward, pull the handle." A browser agent is a sighted employee who's told "open the door" and finds it against any background.

Browserbase: a cloud browser for AI

Browserbase is managed cloud infrastructure for running browsers. Not just headless Chrome on your own server, but a full platform with a few key capabilities:

What Browserbase gives you:

  1. Managed browsers: no need to set up Chromium, versions or dependencies. Browserbase runs the browsers on its infrastructure and you connect over WebSocket.

  2. Session Replay: every agent session can be played back like a video. Indispensable for debugging: you see literally what the agent saw at the moment of the error.

  3. Stealth mode: the browsers are configured to look like a regular user, with proper headers and a realistic browser fingerprint. This often gets past basic anti-bot checks, but not always past serious protection.

  4. Parallelism: you can run many sessions at once (the limit depends on your plan). Monitoring dozens of sites at the same time is an ordinary task.

  5. Residential proxies: the option to go out through real home IP addresses in different countries.

Pricing: Browserbase has a free plan and paid plans that differ in browser hours per month and the number of browsers running at once. The free plan is enough to test an idea. For current plans and prices, see browserbase.com/pricing.

Stagehand: AI-powered Playwright

Stagehand is an open-source SDK from the Browserbase team that adds AI intelligence to browser automation (originally on top of Playwright). If Playwright is the hands (the tools for controlling the browser), Stagehand is the brain that decides what the hands should do.

About versions. In version 3, the act(), extract() and observe() methods are called on the stagehand object itself, and you get the page from stagehand.context.pages()[0]. In version 2 they were called on page (page.act(...)). The examples below are written for version 3; before you start, check the documentation for the version you have installed: docs.stagehand.dev.

Three key methods:

act(instruction): perform an action

typescript
await stagehand.act('Click the "Sign in" button');
await stagehand.act('Fill in the email field with [email protected]');
await stagehand.act('Scroll down the page to the "Pricing" section');

Stagehand takes your plain-language instruction, looks at the page through a screenshot or the DOM, figures out which element is needed and performs the action. If the element has moved, no problem: the AI will find it again.

extract(instruction, schema): pull out data

typescript
const products = await stagehand.extract(
  'Find all products with prices on the page',
  z.object({
    items: z.array(z.object({
      name: z.string(),
      price: z.number(),
      inStock: z.boolean(),
    }))
  })
);

The method returns structured data extracted from the page (validated with Zod). The AI understands what a "price" and "in stock" are, even when different sites show them differently.

observe(instruction): look and report

typescript
const elements = await stagehand.observe(
  'Which buttons are available for actions on this product?'
);
// returns a list of the elements found and suggested actions,
// for example the "Add to cart", "Buy now" and "Compare" buttons

Observing without acting is useful for making decisions inside an agent loop.

Full example: monitoring a competitor's prices

typescript
import { Stagehand } from '@browserbasehq/stagehand';
import { z } from 'zod';

const PriceSchema = z.object({
  products: z.array(z.object({
    name: z.string().describe('Product name'),
    price: z.number().describe('Price in dollars'),
    oldPrice: z.number().nullable().describe('Old price before the discount'),
    inStock: z.boolean().describe('Whether it is in stock'),
    url: z.string().describe('Link to the product'),
  }))
});

async function monitorCompetitorPrices(competitorUrl: string) {
  // Browserbase and model keys are read from .env / environment variables
  const stagehand = new Stagehand({
    env: 'BROWSERBASE',
    model: 'anthropic/claude-sonnet-5-5',   // current models: see the What's current page
    verbose: 1,
  });

  await stagehand.init();
  const page = stagehand.context.pages()[0];

  try {
    // Go to the catalog page
    await page.goto(competitorUrl);
    await page.waitForLoadState('networkidle');

    // If there's a cookie prompt, handle it
    const cookieBanner = await stagehand.observe(
      'Is there a cookie or GDPR banner that needs to be closed?'
    );
    if (cookieBanner.length > 0) {
      await stagehand.act('Click "Accept all" or close the cookie banner');
    }

    // Extract the products from the current page
    const result = await stagehand.extract(
      'Find all products in the catalog: name, price, old price and availability',
      PriceSchema
    );

    // Check whether there's pagination
    const hasNextPage = await stagehand.observe(
      'Is there a "Next page" or "Load more" button?'
    );

    if (hasNextPage.length > 0) {
      await stagehand.act('Click to go to the next page');
      await page.waitForLoadState('networkidle');
      
      const nextPageResult = await stagehand.extract(
        'Find all products on this page',
        PriceSchema
      );
      
      result.products.push(...nextPageResult.products);
    }

    return result.products;

  } finally {
    await stagehand.close();
  }
}

// Run
const prices = await monitorCompetitorPrices('https://competitor.com/catalog/laptops');
console.log(`Found ${prices.length} products`);

Compared with Playwright MCP

In the Browser automation lesson we worked with Playwright and Claude Code, a tool that lets you control a browser as part of a conversation with Claude. It's a great tool for interactive automation guided by a person.

Feature Playwright MCP Stagehand + Browserbase
Control Interactive (through Claude Code) Autonomous (the agent decides)
Infrastructure Local browser Cloud (many in parallel)
Anti-detection Basic Advanced (stealth + proxies, no guarantees)
Session replay No Yes
Scaling 1 browser Many at once (depends on plan)
Best for One-off tasks with a person involved Recurring autonomous tasks

For a production agent that runs on a schedule and handles dozens of sites, Stagehand is the choice. For exploring a site interactively in Claude Code, Playwright MCP is more convenient.

DIY approach: Puppeteer + Claude API

If you don't want to pay for Browserbase, you can build something similar yourself. The stealth plugin for puppeteer-extra isn't updated regularly, so check that it still works before you rely on it:

typescript
import puppeteer from 'puppeteer-extra';
import StealthPlugin from 'puppeteer-extra-plugin-stealth';
import Anthropic from '@anthropic-ai/sdk';

puppeteer.use(StealthPlugin());

const client = new Anthropic();

async function aiScrapePage(url: string, task: string) {
  const browser = await puppeteer.launch({ headless: true });
  const page = await browser.newPage();
  
  // Imitate a real browser
  await page.setUserAgent(
    'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36'
  );
  
  await page.goto(url, { waitUntil: 'networkidle2' });
  
  // Take a screenshot and pass it to Claude
  const screenshot = await page.screenshot({ encoding: 'base64' });
  const htmlContent = await page.content();
  
  const response = await client.messages.create({
    model: 'claude-sonnet-5-5',   // current models: see the What's current page
    max_tokens: 2000,
    messages: [{
      role: 'user',
      content: [
        {
          type: 'image',
          source: { type: 'base64', media_type: 'image/png', data: screenshot as string }
        },
        {
          type: 'text',
          text: `Task: ${task}\n\nPage HTML:\n${htmlContent.substring(0, 5000)}\n\nReturn the result as JSON.`
        }
      ]
    }]
  });
  
  await browser.close();
  return response.content.map((b) => (b.type === 'text' ? b.text : '')).join('');
}

This approach costs nothing for infrastructure (you only pay for the Claude API), but it doesn't scale as well and takes more setup.

Anti-detection: how not to get banned

Sites use several layers of bot protection:

Level 1, basic (easy to get past):

  • User-Agent checks: solved by setting a real UA
  • Requests that come too fast: solved with random pauses (Math.random() * 2000 + 500 ms)
  • No cookies/localStorage: solved by using a normal browser

Level 2, advanced (takes effort):

  • Browser fingerprinting (canvas, WebGL, fonts): Browserbase and puppeteer-stealth partly hide it
  • Mouse behavior analysis: you need realistic movements
  • IP reputation: residential proxies from providers like Bright Data or Oxylabs

Level 3, serious protection:

  • CAPTCHA (reCAPTCHA v3, hCaptcha): there are paid solving services like 2captcha.com or CapSolver (check their sites for prices)
  • Cloudflare Bot Management: no service can guarantee getting past it

Important. Getting around bot protection often violates a site's terms of use, and in some countries the law as well. This lesson shows the mechanics so you understand how it works. If a site has an official API, or you can arrange access with the owner, that's more reliable and safer.

typescript
// A schematic example of a 2captcha integration (check the service's API and the site's terms)
import Captcha2 from '2captcha-ts';
const solver = new Captcha2.Solver(process.env.CAPTCHA_API_KEY);

// When a CAPTCHA is detected
const siteKey = await page.evaluate(() => 
  document.querySelector('[data-sitekey]')?.getAttribute('data-sitekey')
);

if (siteKey) {
  const result = await solver.recaptcha({
    pageurl: page.url(),
    googlekey: siteKey,
  });
  
  await page.evaluate((token) => {
    (window as any).grecaptcha?.enterprise?.execute?.(token);
  }, result.data);
}

Ethics and legality: what's OK and what isn't

🎨 Picture this: scraping is like visiting a public library. The books are open to everyone: you can read them, take notes, analyze them. But you can't tear out pages, photograph the closed archives or pretend to be on staff.

What's usually OK (public data):

  • ✅ Prices on stores' public pages
  • ✅ Publicly available listings (real estate, jobs)
  • ✅ News articles (for aggregation, not republishing)
  • ✅ Data from public APIs
  • ✅ Information about companies from open sources

What isn't (legal risk):

  • ❌ PII: users' personal data (emails, phone numbers from restricted sections)
  • ❌ Paywalled content without a subscription
  • ❌ Data whose scraping is explicitly prohibited in the Terms of Service
  • ❌ Load that interferes with how the site works (DDoS-like behavior)
  • ❌ Getting around authentication systems

Before you launch, always check the target site's robots.txt and Terms of Service. In the EU, also look at GDPR; in the US, the CFAA and the scraping rulings in LinkedIn vs. hiQ.

Rule of thumb: if the information is shown to any visitor who isn't logged in and hasn't registered, scraping that public data is most likely acceptable. If you need a login, you're already in a gray area.

Monetization: selling data feeds

Browser agents aren't just a tool for your own needs. They're a product you can sell.

Ways to sell:

  1. A regular data feed: the client pays for a periodic export. For example, a daily dump of competitor prices in a particular niche. You set up the agent once, then it runs on a schedule while you watch for errors and fix the script when the site changes.

  2. A one-off data collection project: the client pays for a single export. A list of companies in a particular region, a list of potential partners, a market analysis.

  3. Monitoring as a service: tracking changes on websites (new job postings, price changes, new competitors showing up). The client gets alerts by email, Slack or a messenger app.

  4. A white-label solution: you sell the tool, not the data. You set up the agent for the client and they use it themselves. A one-time fee plus support.

The price is set by the market and your costs (servers, your Browserbase plan, tokens, time spent on support). Income in this niche isn't guaranteed: estimate your costs and the demand in your market using the Packaging your offer and Pricing and monetization lessons.

🎨 Picture this: a farmer has installed automatic irrigation on several fields. The fields run on their own; the farmer just checks on them. A browser agent is your irrigation system, and the data is the harvest a client is willing to pay for.


Practice

Task: a price monitoring agent with chat alerts

We'll build an agent that checks a competitor's product prices once an hour and sends an alert if a price drops by more than 5%. The example sends alerts through a Telegram bot because it takes only a token and a chat ID to set up; the same sendAlert function can just as well send an email or a Slack message.

Step 1: Install the dependencies

bash
mkdir price-monitor-agent && cd price-monitor-agent
npm init -y
npm install @browserbasehq/stagehand zod dotenv node-telegram-bot-api
npm install -D typescript @types/node tsx

Create .env:

env
BROWSERBASE_API_KEY=bb_live_xxxx
BROWSERBASE_PROJECT_ID=prj_xxxx
TELEGRAM_BOT_TOKEN=xxxx
TELEGRAM_CHAT_ID=xxxx
ANTHROPIC_API_KEY=sk-ant-xxxx

Step 2: Data schema and configuration

Create src/types.ts:

typescript
import { z } from 'zod';

export const ProductSchema = z.object({
  name: z.string(),
  price: z.number(),
  oldPrice: z.number().nullable(),
  inStock: z.boolean(),
  sku: z.string().optional(),
});

export type Product = z.infer<typeof ProductSchema>;

export const MonitorConfig = {
  targetUrl: 'https://example-shop.com/catalog/laptops',
  checkIntervalMinutes: 60,
  priceDropThresholdPercent: 5,
};

Step 3: The scraping function

Create src/scraper.ts:

typescript
import { Stagehand } from '@browserbasehq/stagehand';
import { z } from 'zod';
import { ProductSchema, type Product } from './types';

export async function scrapeProducts(url: string): Promise<Product[]> {
  // Browserbase and model keys are read from .env (dotenv loads the .env file)
  const stagehand = new Stagehand({
    env: 'BROWSERBASE',
    model: 'anthropic/claude-sonnet-5-5',
  });

  await stagehand.init();
  const page = stagehand.context.pages()[0];

  try {
    await page.goto(url);
    await page.waitForLoadState('networkidle');

    // Handle cookie banners
    const cookieBanner = await stagehand.observe(
      'Is there a cookie or consent banner that needs to be accepted?'
    );
    if (cookieBanner.length > 0) {
      await stagehand.act('Accept or close the cookie banner');
      await page.waitForTimeout(1000);
    }

    // Extract the product data
    const result = await stagehand.extract(
      `
        Find all products on the catalog page.
        For each one, extract: the exact name, the current price as a number,
        the old price (if there's a discount), and whether it's in stock.
        Prices must be numbers without currency symbols.
      `,
      z.object({
        products: z.array(ProductSchema)
      })
    );

    return result.products;

  } finally {
    await stagehand.close();
  }
}

Step 4: Storing and comparing prices

Create src/storage.ts:

typescript
import fs from 'fs/promises';
import path from 'path';
import type { Product } from './types';

const DATA_FILE = path.join(process.cwd(), 'prices.json');

export interface PriceRecord {
  timestamp: string;
  products: Product[];
}

export async function saveProducts(products: Product[]): Promise<void> {
  const record: PriceRecord = {
    timestamp: new Date().toISOString(),
    products,
  };
  await fs.writeFile(DATA_FILE, JSON.stringify(record, null, 2));
}

export async function loadPreviousProducts(): Promise<Product[] | null> {
  try {
    const data = await fs.readFile(DATA_FILE, 'utf-8');
    const record: PriceRecord = JSON.parse(data);
    return record.products;
  } catch {
    return null; // First run
  }
}

export interface PriceDrop {
  product: Product;
  oldPrice: number;
  newPrice: number;
  dropPercent: number;
}

export function findPriceDrops(
  current: Product[],
  previous: Product[],
  thresholdPercent: number
): PriceDrop[] {
  const drops: PriceDrop[] = [];
  
  for (const currentProduct of current) {
    const prevProduct = previous.find(p => p.name === currentProduct.name);
    if (!prevProduct) continue;
    
    const dropPercent = ((prevProduct.price - currentProduct.price) / prevProduct.price) * 100;
    
    if (dropPercent >= thresholdPercent) {
      drops.push({
        product: currentProduct,
        oldPrice: prevProduct.price,
        newPrice: currentProduct.price,
        dropPercent,
      });
    }
  }
  
  return drops.sort((a, b) => b.dropPercent - a.dropPercent);
}

Step 5: Alerts and the main loop

Create src/index.ts:

typescript
import 'dotenv/config';
import TelegramBot from 'node-telegram-bot-api';
import { scrapeProducts } from './scraper';
import { saveProducts, loadPreviousProducts, findPriceDrops } from './storage';
import { MonitorConfig } from './types';

const bot = new TelegramBot(process.env.TELEGRAM_BOT_TOKEN!);

async function sendAlert(drops: ReturnType<typeof findPriceDrops>) {
  const chatId = process.env.TELEGRAM_CHAT_ID!;
  
  let message = `🔥 *Price drop detected!*\n\n`;
  
  for (const drop of drops) {
    message += `📦 *${drop.product.name}*\n`;
    message += `💰 $${drop.oldPrice.toLocaleString()} → $${drop.newPrice.toLocaleString()}\n`;
    message += `📉 Drop: *${drop.dropPercent.toFixed(1)}%*\n`;
    message += drop.product.inStock ? '✅ In stock\n' : '❌ Out of stock\n';
    message += '\n';
  }
  
  await bot.sendMessage(chatId, message, { parse_mode: 'Markdown' });
}

async function runCheck() {
  console.log(`[${new Date().toISOString()}] Starting price check...`);
  
  try {
    // Scrape the current prices
    const currentProducts = await scrapeProducts(MonitorConfig.targetUrl);
    console.log(`Products found: ${currentProducts.length}`);
    
    // Compare with the previous data
    const previousProducts = await loadPreviousProducts();
    
    if (previousProducts) {
      const drops = findPriceDrops(
        currentProducts,
        previousProducts,
        MonitorConfig.priceDropThresholdPercent
      );
      
      if (drops.length > 0) {
        console.log(`Price drops found: ${drops.length}`);
        await sendAlert(drops);
      } else {
        console.log('No significant price changes detected');
      }
    } else {
      console.log('First run: baseline prices saved');
    }
    
    // Save the current prices as the baseline
    await saveProducts(currentProducts);
    
  } catch (error) {
    console.error('Error during check:', error);
    await bot.sendMessage(
      process.env.TELEGRAM_CHAT_ID!,
      `⚠️ Monitoring agent error: ${error}`
    );
  }
}

// Run now and repeat on a schedule
async function main() {
  console.log('Monitoring agent started');
  
  // First run right away
  await runCheck();
  
  // Then on a schedule
  setInterval(
    runCheck,
    MonitorConfig.checkIntervalMinutes * 60 * 1000
  );
}

main().catch(console.error);

Run it:

bash
npx tsx src/index.ts

For production: a simple VPS with pm2 start, or a single pass run on a schedule with cron instead of setInterval.


Tools and resources

Tool Purpose Link
Browserbase Managed cloud browsers for agents browserbase.com
Stagehand AI-powered Playwright SDK (open source) github.com/browserbase/stagehand
Playwright The underlying browser automation engine playwright.dev
Puppeteer + Stealth DIY alternative with an anti-detection plugin github.com/berstend/puppeteer-extra
2captcha CAPTCHA-solving service (price on their site) 2captcha.com
CapSolver An AI-based alternative to 2captcha capsolver.com
Bright Data Residential proxies to get past blocks brightdata.com

Key takeaways

"A browser agent isn't an improved script. It's an employee who understands the goal instead of just following instructions. When a site changes, a person adapts. So does the agent."

"A monitoring agent you've set up once can be packaged as a service, but it isn't passive income: sites change and the agent needs upkeep. What it brings in depends on the niche and the clients, with no guarantees. Browserbase and Stagehand let you put together a working prototype quickly."

"Public data is like a shop window. Looking and analyzing is your right. But breaking the lock or copying someone else's content for commercial use is another story. Know where the line is."


Next lesson

→ Multi-agent orchestration: LangGraph, CrewAI and Mastra

If one agent is one employee, a multi-agent system is a whole department. We'll look at how agents coordinate, hand tasks off to each other, and how to build a pipeline out of dozens of specialized AI workers. LangGraph for stateful workflows, CrewAI for role-based teams, and Mastra as an open-source TypeScript framework.

The mark stays in this browser only and is never sent anywhere. My progress