The gist
That's a browser agent: an AI system that drives a real browser, reads pages the way a person does and decides on its next action based on what it sees. Not just a script with fixed CSS selectors, but real intelligence that understands what a page means.
In this lesson we'll look at how to build agents like this with Browserbase and Stagehand, see how they differ from classic automation, and write a working agent that monitors competitor prices.
Key concepts
- Browser agent vs. classic automation: AI understands what a page means instead of just looking for CSS selectors
- Browserbase: cloud browser infrastructure for AI agents, with anti-detection features
- Stagehand: an SDK from Browserbase built on top of Playwright, with LLM-driven methods
act(),extract()andobserve() - Session replay: recording and replaying sessions to debug agents
- Anti-detection techniques: stealth mode, residential proxies, getting past CAPTCHAs
- Scraping ethics: public vs. private data, Terms of Service, PII
- Monetization: data collection and monitoring services (a service format, not a guaranteed income)
Theory
Why browser agents aren't Selenium 2.0
When Selenium appeared in 2004, web automation was built on one principle: a programmer manually finds the CSS selector or XPath to the element they need and tells the script, "click #buy-button." That worked until a site changed its layout. As soon as the developers renamed the button's class, the script broke.
Puppeteer (2017) and Playwright (2020) made this approach faster and more reliable, but the principle stayed the same: rigid instructions based on the DOM structure. If the page changes, the automation breaks.
A browser agent thinks differently. Instead of "click the element with id=submit," it gets the task "place the order" and figures out on its own what that takes. It reads the text on the page, understands the context, sees that the button says "Place order" or "Checkout," and clicks it. If the layout changes, the agent simply reworks its plan.
Browserbase: a cloud browser for AI
Browserbase is managed cloud infrastructure for running browsers. Not just headless Chrome on your own server, but a full platform with a few key capabilities:
What Browserbase gives you:
Managed browsers: no need to set up Chromium, versions or dependencies. Browserbase runs the browsers on its infrastructure and you connect over WebSocket.
Session Replay: every agent session can be played back like a video. Indispensable for debugging: you see literally what the agent saw at the moment of the error.
Stealth mode: the browsers are configured to look like a regular user, with proper headers and a realistic browser fingerprint. This often gets past basic anti-bot checks, but not always past serious protection.
Parallelism: you can run many sessions at once (the limit depends on your plan). Monitoring dozens of sites at the same time is an ordinary task.
Residential proxies: the option to go out through real home IP addresses in different countries.
Pricing: Browserbase has a free plan and paid plans that differ in browser hours per month and the number of browsers running at once. The free plan is enough to test an idea. For current plans and prices, see browserbase.com/pricing.
Stagehand: AI-powered Playwright
Stagehand is an open-source SDK from the Browserbase team that adds AI intelligence to browser automation (originally on top of Playwright). If Playwright is the hands (the tools for controlling the browser), Stagehand is the brain that decides what the hands should do.
About versions. In version 3, the act(), extract() and observe() methods are called on the stagehand object itself, and you get the page from stagehand.context.pages()[0]. In version 2 they were called on page (page.act(...)). The examples below are written for version 3; before you start, check the documentation for the version you have installed: docs.stagehand.dev.
Three key methods:
act(instruction): perform an action
await stagehand.act('Click the "Sign in" button');
await stagehand.act('Fill in the email field with [email protected]');
await stagehand.act('Scroll down the page to the "Pricing" section');Stagehand takes your plain-language instruction, looks at the page through a screenshot or the DOM, figures out which element is needed and performs the action. If the element has moved, no problem: the AI will find it again.
extract(instruction, schema): pull out data
const products = await stagehand.extract(
'Find all products with prices on the page',
z.object({
items: z.array(z.object({
name: z.string(),
price: z.number(),
inStock: z.boolean(),
}))
})
);The method returns structured data extracted from the page (validated with Zod). The AI understands what a "price" and "in stock" are, even when different sites show them differently.
observe(instruction): look and report
const elements = await stagehand.observe(
'Which buttons are available for actions on this product?'
);
// returns a list of the elements found and suggested actions,
// for example the "Add to cart", "Buy now" and "Compare" buttonsObserving without acting is useful for making decisions inside an agent loop.
Full example: monitoring a competitor's prices
import { Stagehand } from '@browserbasehq/stagehand';
import { z } from 'zod';
const PriceSchema = z.object({
products: z.array(z.object({
name: z.string().describe('Product name'),
price: z.number().describe('Price in dollars'),
oldPrice: z.number().nullable().describe('Old price before the discount'),
inStock: z.boolean().describe('Whether it is in stock'),
url: z.string().describe('Link to the product'),
}))
});
async function monitorCompetitorPrices(competitorUrl: string) {
// Browserbase and model keys are read from .env / environment variables
const stagehand = new Stagehand({
env: 'BROWSERBASE',
model: 'anthropic/claude-sonnet-5-5', // current models: see the What's current page
verbose: 1,
});
await stagehand.init();
const page = stagehand.context.pages()[0];
try {
// Go to the catalog page
await page.goto(competitorUrl);
await page.waitForLoadState('networkidle');
// If there's a cookie prompt, handle it
const cookieBanner = await stagehand.observe(
'Is there a cookie or GDPR banner that needs to be closed?'
);
if (cookieBanner.length > 0) {
await stagehand.act('Click "Accept all" or close the cookie banner');
}
// Extract the products from the current page
const result = await stagehand.extract(
'Find all products in the catalog: name, price, old price and availability',
PriceSchema
);
// Check whether there's pagination
const hasNextPage = await stagehand.observe(
'Is there a "Next page" or "Load more" button?'
);
if (hasNextPage.length > 0) {
await stagehand.act('Click to go to the next page');
await page.waitForLoadState('networkidle');
const nextPageResult = await stagehand.extract(
'Find all products on this page',
PriceSchema
);
result.products.push(...nextPageResult.products);
}
return result.products;
} finally {
await stagehand.close();
}
}
// Run
const prices = await monitorCompetitorPrices('https://competitor.com/catalog/laptops');
console.log(`Found ${prices.length} products`);Compared with Playwright MCP
In the Browser automation lesson we worked with Playwright and Claude Code, a tool that lets you control a browser as part of a conversation with Claude. It's a great tool for interactive automation guided by a person.
| Feature | Playwright MCP | Stagehand + Browserbase |
|---|---|---|
| Control | Interactive (through Claude Code) | Autonomous (the agent decides) |
| Infrastructure | Local browser | Cloud (many in parallel) |
| Anti-detection | Basic | Advanced (stealth + proxies, no guarantees) |
| Session replay | No | Yes |
| Scaling | 1 browser | Many at once (depends on plan) |
| Best for | One-off tasks with a person involved | Recurring autonomous tasks |
For a production agent that runs on a schedule and handles dozens of sites, Stagehand is the choice. For exploring a site interactively in Claude Code, Playwright MCP is more convenient.
DIY approach: Puppeteer + Claude API
If you don't want to pay for Browserbase, you can build something similar yourself. The stealth plugin for puppeteer-extra isn't updated regularly, so check that it still works before you rely on it:
import puppeteer from 'puppeteer-extra';
import StealthPlugin from 'puppeteer-extra-plugin-stealth';
import Anthropic from '@anthropic-ai/sdk';
puppeteer.use(StealthPlugin());
const client = new Anthropic();
async function aiScrapePage(url: string, task: string) {
const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
// Imitate a real browser
await page.setUserAgent(
'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36'
);
await page.goto(url, { waitUntil: 'networkidle2' });
// Take a screenshot and pass it to Claude
const screenshot = await page.screenshot({ encoding: 'base64' });
const htmlContent = await page.content();
const response = await client.messages.create({
model: 'claude-sonnet-5-5', // current models: see the What's current page
max_tokens: 2000,
messages: [{
role: 'user',
content: [
{
type: 'image',
source: { type: 'base64', media_type: 'image/png', data: screenshot as string }
},
{
type: 'text',
text: `Task: ${task}\n\nPage HTML:\n${htmlContent.substring(0, 5000)}\n\nReturn the result as JSON.`
}
]
}]
});
await browser.close();
return response.content.map((b) => (b.type === 'text' ? b.text : '')).join('');
}This approach costs nothing for infrastructure (you only pay for the Claude API), but it doesn't scale as well and takes more setup.
Anti-detection: how not to get banned
Sites use several layers of bot protection:
Level 1, basic (easy to get past):
- User-Agent checks: solved by setting a real UA
- Requests that come too fast: solved with random pauses (
Math.random() * 2000 + 500ms) - No cookies/localStorage: solved by using a normal browser
Level 2, advanced (takes effort):
- Browser fingerprinting (canvas, WebGL, fonts): Browserbase and puppeteer-stealth partly hide it
- Mouse behavior analysis: you need realistic movements
- IP reputation: residential proxies from providers like Bright Data or Oxylabs
Level 3, serious protection:
- CAPTCHA (reCAPTCHA v3, hCaptcha): there are paid solving services like 2captcha.com or CapSolver (check their sites for prices)
- Cloudflare Bot Management: no service can guarantee getting past it
Important. Getting around bot protection often violates a site's terms of use, and in some countries the law as well. This lesson shows the mechanics so you understand how it works. If a site has an official API, or you can arrange access with the owner, that's more reliable and safer.
// A schematic example of a 2captcha integration (check the service's API and the site's terms)
import Captcha2 from '2captcha-ts';
const solver = new Captcha2.Solver(process.env.CAPTCHA_API_KEY);
// When a CAPTCHA is detected
const siteKey = await page.evaluate(() =>
document.querySelector('[data-sitekey]')?.getAttribute('data-sitekey')
);
if (siteKey) {
const result = await solver.recaptcha({
pageurl: page.url(),
googlekey: siteKey,
});
await page.evaluate((token) => {
(window as any).grecaptcha?.enterprise?.execute?.(token);
}, result.data);
}Ethics and legality: what's OK and what isn't
What's usually OK (public data):
- ✅ Prices on stores' public pages
- ✅ Publicly available listings (real estate, jobs)
- ✅ News articles (for aggregation, not republishing)
- ✅ Data from public APIs
- ✅ Information about companies from open sources
What isn't (legal risk):
- ❌ PII: users' personal data (emails, phone numbers from restricted sections)
- ❌ Paywalled content without a subscription
- ❌ Data whose scraping is explicitly prohibited in the Terms of Service
- ❌ Load that interferes with how the site works (DDoS-like behavior)
- ❌ Getting around authentication systems
Before you launch, always check the target site's robots.txt and Terms of Service. In the EU, also look at GDPR; in the US, the CFAA and the scraping rulings in LinkedIn vs. hiQ.
Rule of thumb: if the information is shown to any visitor who isn't logged in and hasn't registered, scraping that public data is most likely acceptable. If you need a login, you're already in a gray area.
Monetization: selling data feeds
Browser agents aren't just a tool for your own needs. They're a product you can sell.
Ways to sell:
A regular data feed: the client pays for a periodic export. For example, a daily dump of competitor prices in a particular niche. You set up the agent once, then it runs on a schedule while you watch for errors and fix the script when the site changes.
A one-off data collection project: the client pays for a single export. A list of companies in a particular region, a list of potential partners, a market analysis.
Monitoring as a service: tracking changes on websites (new job postings, price changes, new competitors showing up). The client gets alerts by email, Slack or a messenger app.
A white-label solution: you sell the tool, not the data. You set up the agent for the client and they use it themselves. A one-time fee plus support.
The price is set by the market and your costs (servers, your Browserbase plan, tokens, time spent on support). Income in this niche isn't guaranteed: estimate your costs and the demand in your market using the Packaging your offer and Pricing and monetization lessons.
Practice
Task: a price monitoring agent with chat alerts
We'll build an agent that checks a competitor's product prices once an hour and sends an alert if a price drops by more than 5%. The example sends alerts through a Telegram bot because it takes only a token and a chat ID to set up; the same sendAlert function can just as well send an email or a Slack message.
Step 1: Install the dependencies
mkdir price-monitor-agent && cd price-monitor-agent
npm init -y
npm install @browserbasehq/stagehand zod dotenv node-telegram-bot-api
npm install -D typescript @types/node tsxCreate .env:
BROWSERBASE_API_KEY=bb_live_xxxx
BROWSERBASE_PROJECT_ID=prj_xxxx
TELEGRAM_BOT_TOKEN=xxxx
TELEGRAM_CHAT_ID=xxxx
ANTHROPIC_API_KEY=sk-ant-xxxxStep 2: Data schema and configuration
Create src/types.ts:
import { z } from 'zod';
export const ProductSchema = z.object({
name: z.string(),
price: z.number(),
oldPrice: z.number().nullable(),
inStock: z.boolean(),
sku: z.string().optional(),
});
export type Product = z.infer<typeof ProductSchema>;
export const MonitorConfig = {
targetUrl: 'https://example-shop.com/catalog/laptops',
checkIntervalMinutes: 60,
priceDropThresholdPercent: 5,
};Step 3: The scraping function
Create src/scraper.ts:
import { Stagehand } from '@browserbasehq/stagehand';
import { z } from 'zod';
import { ProductSchema, type Product } from './types';
export async function scrapeProducts(url: string): Promise<Product[]> {
// Browserbase and model keys are read from .env (dotenv loads the .env file)
const stagehand = new Stagehand({
env: 'BROWSERBASE',
model: 'anthropic/claude-sonnet-5-5',
});
await stagehand.init();
const page = stagehand.context.pages()[0];
try {
await page.goto(url);
await page.waitForLoadState('networkidle');
// Handle cookie banners
const cookieBanner = await stagehand.observe(
'Is there a cookie or consent banner that needs to be accepted?'
);
if (cookieBanner.length > 0) {
await stagehand.act('Accept or close the cookie banner');
await page.waitForTimeout(1000);
}
// Extract the product data
const result = await stagehand.extract(
`
Find all products on the catalog page.
For each one, extract: the exact name, the current price as a number,
the old price (if there's a discount), and whether it's in stock.
Prices must be numbers without currency symbols.
`,
z.object({
products: z.array(ProductSchema)
})
);
return result.products;
} finally {
await stagehand.close();
}
}Step 4: Storing and comparing prices
Create src/storage.ts:
import fs from 'fs/promises';
import path from 'path';
import type { Product } from './types';
const DATA_FILE = path.join(process.cwd(), 'prices.json');
export interface PriceRecord {
timestamp: string;
products: Product[];
}
export async function saveProducts(products: Product[]): Promise<void> {
const record: PriceRecord = {
timestamp: new Date().toISOString(),
products,
};
await fs.writeFile(DATA_FILE, JSON.stringify(record, null, 2));
}
export async function loadPreviousProducts(): Promise<Product[] | null> {
try {
const data = await fs.readFile(DATA_FILE, 'utf-8');
const record: PriceRecord = JSON.parse(data);
return record.products;
} catch {
return null; // First run
}
}
export interface PriceDrop {
product: Product;
oldPrice: number;
newPrice: number;
dropPercent: number;
}
export function findPriceDrops(
current: Product[],
previous: Product[],
thresholdPercent: number
): PriceDrop[] {
const drops: PriceDrop[] = [];
for (const currentProduct of current) {
const prevProduct = previous.find(p => p.name === currentProduct.name);
if (!prevProduct) continue;
const dropPercent = ((prevProduct.price - currentProduct.price) / prevProduct.price) * 100;
if (dropPercent >= thresholdPercent) {
drops.push({
product: currentProduct,
oldPrice: prevProduct.price,
newPrice: currentProduct.price,
dropPercent,
});
}
}
return drops.sort((a, b) => b.dropPercent - a.dropPercent);
}Step 5: Alerts and the main loop
Create src/index.ts:
import 'dotenv/config';
import TelegramBot from 'node-telegram-bot-api';
import { scrapeProducts } from './scraper';
import { saveProducts, loadPreviousProducts, findPriceDrops } from './storage';
import { MonitorConfig } from './types';
const bot = new TelegramBot(process.env.TELEGRAM_BOT_TOKEN!);
async function sendAlert(drops: ReturnType<typeof findPriceDrops>) {
const chatId = process.env.TELEGRAM_CHAT_ID!;
let message = `🔥 *Price drop detected!*\n\n`;
for (const drop of drops) {
message += `📦 *${drop.product.name}*\n`;
message += `💰 $${drop.oldPrice.toLocaleString()} → $${drop.newPrice.toLocaleString()}\n`;
message += `📉 Drop: *${drop.dropPercent.toFixed(1)}%*\n`;
message += drop.product.inStock ? '✅ In stock\n' : '❌ Out of stock\n';
message += '\n';
}
await bot.sendMessage(chatId, message, { parse_mode: 'Markdown' });
}
async function runCheck() {
console.log(`[${new Date().toISOString()}] Starting price check...`);
try {
// Scrape the current prices
const currentProducts = await scrapeProducts(MonitorConfig.targetUrl);
console.log(`Products found: ${currentProducts.length}`);
// Compare with the previous data
const previousProducts = await loadPreviousProducts();
if (previousProducts) {
const drops = findPriceDrops(
currentProducts,
previousProducts,
MonitorConfig.priceDropThresholdPercent
);
if (drops.length > 0) {
console.log(`Price drops found: ${drops.length}`);
await sendAlert(drops);
} else {
console.log('No significant price changes detected');
}
} else {
console.log('First run: baseline prices saved');
}
// Save the current prices as the baseline
await saveProducts(currentProducts);
} catch (error) {
console.error('Error during check:', error);
await bot.sendMessage(
process.env.TELEGRAM_CHAT_ID!,
`⚠️ Monitoring agent error: ${error}`
);
}
}
// Run now and repeat on a schedule
async function main() {
console.log('Monitoring agent started');
// First run right away
await runCheck();
// Then on a schedule
setInterval(
runCheck,
MonitorConfig.checkIntervalMinutes * 60 * 1000
);
}
main().catch(console.error);Run it:
npx tsx src/index.tsFor production: a simple VPS with pm2 start, or a single pass run on a schedule with cron instead of setInterval.
Tools and resources
| Tool | Purpose | Link |
|---|---|---|
| Browserbase | Managed cloud browsers for agents | browserbase.com |
| Stagehand | AI-powered Playwright SDK (open source) | github.com/browserbase/stagehand |
| Playwright | The underlying browser automation engine | playwright.dev |
| Puppeteer + Stealth | DIY alternative with an anti-detection plugin | github.com/berstend/puppeteer-extra |
| 2captcha | CAPTCHA-solving service (price on their site) | 2captcha.com |
| CapSolver | An AI-based alternative to 2captcha | capsolver.com |
| Bright Data | Residential proxies to get past blocks | brightdata.com |
Key takeaways
"A browser agent isn't an improved script. It's an employee who understands the goal instead of just following instructions. When a site changes, a person adapts. So does the agent."
"A monitoring agent you've set up once can be packaged as a service, but it isn't passive income: sites change and the agent needs upkeep. What it brings in depends on the niche and the clients, with no guarantees. Browserbase and Stagehand let you put together a working prototype quickly."
"Public data is like a shop window. Looking and analyzing is your right. But breaking the lock or copying someone else's content for commercial use is another story. Know where the line is."
Next lesson
→ Multi-agent orchestration: LangGraph, CrewAI and Mastra
If one agent is one employee, a multi-agent system is a whole department. We'll look at how agents coordinate, hand tasks off to each other, and how to build a pipeline out of dozens of specialized AI workers. LangGraph for stateful workflows, CrewAI for role-based teams, and Mastra as an open-source TypeScript framework.
The mark stays in this browser only and is never sent anywhere. My progress