Library · Marketing, sales and analytics with AI

Finding potential clients: from scraping websites to a CSV file

Builder110 minUpdated: October 2026
65 of 105 in the library

Time: ~20 min theory + 90 min practice


The gist

🎨 Picture this: a Lead Generation System is like the night shift at a factory. While you sleep, the assembly line gathers, checks and packs the finished product. In the morning, there are boxes on your desk with the names, addresses and a short background on each client.

Imagine you could hire someone who, overnight, gathers 500 contacts of potential clients (leads), visits each one's website, reads up on the business, writes a personalized icebreaker (an opening line for a conversation), and has a finished CSV waiting for you in the morning. That's exactly what a Lead Generation System built with Claude Code does. It's full B2B (selling to businesses) automation: from an empty list to an enriched database you can take straight into cold outreach (contacting potential clients you don't know yet).


Key concepts

  • Lead Generation Pipeline (a processing assembly line): 4 stages: collect → filter → enrich → export
  • Apify: a marketplace of scrapers (programs that collect data), including a Google Maps scraper for finding businesses by niche and location
  • Filtering: keep only leads with a website and a phone number (drop the rest)
  • Enrichment: Claude analyzes each website: value proposition, what makes them different, a case study (a write-up of a real project with results)
  • Output: a CSV with the fields company, location, website, phone, value_prop, differentiator, case_study
  • Skill (a reusable Claude module): the whole pipeline gets packaged into a reusable skill with a slash command

Theory

How the pipeline is structured

🎨 Picture this: The pipeline is like an oil pipeline with four sections. Each section refines the raw material: first you extract it (scrape), then filter out impurities (only leads with a website and phone), then refine it (Claude reads each website), then bottle it (CSV). What comes out isn't crude oil, it's gasoline.

Code
Input:
  niche = "roofing companies"
  location = "North Carolina"
  count = 30

Step 1: Google Maps Scraping (Apify)
  → 34 businesses with Google Maps metadata

Step 2: Filtering
  → Keep only: website AND phone exist
  → For example: 22 of 34 pass the filter

Step 3: Website Scraping + Claude Enrichment
  → For each of the 22: scrape homepage + /about + /services
  → Claude extracts: value_prop, differentiator, case_study
  
Step 4: CSV Export
  → leads.csv — ready for Google Sheets and your CRM

Why Apify and not scraping Google directly

🎨 Picture this: Apify is like a professional fishing trawler. You could go out alone with a fishing rod (scrape Google directly), but you'd get caught and kicked out fast. A licensed trawler works the open sea and brings back the full catch, neatly sorted.

Google Maps has an official paid API (application programming interface), but it isn't designed for exporting large business databases. Scraping Google's pages directly violates its ToS (Terms of Service) and gets you blocked quickly. Terms of service change: before you run anything, check the current terms and the laws where you live.

Apify is a marketplace of ready-made scrapers with:

  • Proxy rotation (to avoid getting blocked)
  • Rate limit management
  • Structured output (JSON)
  • A free plan with a small monthly credit (current terms: apify.com/pricing)

The Google Maps scraper on Apify takes a query like "roofing companies near North Carolina" and returns a list of businesses with address, website, phone, rating and hours.


The opening prompt (your request to the AI): explaining the task to Claude Code

An important pattern from practice: describe the whole task first; don't jump straight to "create a skill." Let Claude build and test the system first, then package it into a skill:

Type this into the chat
Hey, today's task is to build a Lead Generation System.

Step 1: I give you a niche, a location and a number of results
Step 2: you go to Apify and use the Google Maps scraper
Step 3: you filter, keeping only leads WITH a website and a phone number
Step 4: you visit each website (homepage + about + services)
Step 5: Claude extracts: value proposition, what makes them unique, client case studies
Step 6: you output a CSV: company_name, location, website, phone, value_prop, differentiator, case_study

Ask clarifying questions before you start.

Claude will ask: do you have an Apify API key, which output format do you prefer, which language to use.


Python script: the system architecture

python
# lead_generation.py
import os
import csv
import json
import anthropic
from apify_client import ApifyClient

APIFY_TOKEN = os.environ["APIFY_API_TOKEN"]
ANTHROPIC_API_KEY = os.environ["ANTHROPIC_API_KEY"]

def scrape_google_maps(niche: str, location: str, count: int) -> list[dict]:
    """Step 1: Collect businesses with the Apify Google Maps scraper"""
    client = ApifyClient(APIFY_TOKEN)
    
    run_input = {
        "searchStringsArray": [f"{niche} near {location}"],
        "maxCrawledPlacesPerSearch": count,
        "language": "en",
        "countryCode": "us",
    }
    
    # Google Maps Scraper on Apify (compass/crawler-google-places)
    run = client.actor("nwua9Gu5YrADL7ZDj").call(run_input=run_input)
    
    results = []
    for item in client.dataset(run["defaultDatasetId"]).iterate_items():
        results.append({
            "company_name": item.get("title", ""),
            "location": item.get("address", ""),
            "website": item.get("website", ""),
            "phone": item.get("phone", ""),
            "rating": item.get("totalScore", ""),
        })
    
    return results


def filter_leads(raw_leads: list[dict]) -> list[dict]:
    """Step 2: Keep only leads with a website and a phone number"""
    filtered = [
        lead for lead in raw_leads
        if lead.get("website") and lead.get("phone")
    ]
    print(f"Filtering: {len(raw_leads)} → {len(filtered)} leads")
    return filtered


def scrape_website(url: str) -> str:
    """Step 3a: Collect text from the website (homepage + about + services)"""
    import requests
    from bs4 import BeautifulSoup
    
    pages_to_scrape = [url, f"{url}/about", f"{url}/services"]
    all_text = []
    
    for page_url in pages_to_scrape:
        try:
            response = requests.get(page_url, timeout=10, headers={
                "User-Agent": "Mozilla/5.0 (compatible; LeadBot/1.0)"
            })
            if response.status_code == 200:
                soup = BeautifulSoup(response.text, "html.parser")
                # Remove scripts and styles
                for tag in soup(["script", "style", "nav", "footer"]):
                    tag.decompose()
                text = soup.get_text(separator=" ", strip=True)
                all_text.append(text[:2000])  # Take the first 2000 characters of each page
        except Exception:
            continue
    
    return " | ".join(all_text)


def enrich_with_claude(company_name: str, website_text: str) -> dict:
    """Step 3b: Claude extracts structured insights from the website text"""
    client = anthropic.Anthropic(api_key=ANTHROPIC_API_KEY)
    
    message = client.messages.create(
        model="claude-haiku-4-5",  # Haiku is enough for structured extraction; current models: the What's current page
        max_tokens=500,
        messages=[{
            "role": "user",
            "content": f"""Analyze this text from the website of {company_name}.

Website text:
{website_text[:3000]}

Extract as JSON:
- value_prop: the main value proposition (1-2 sentences)
- differentiator: what makes them unique among competitors (1 sentence)
- case_study: the client's name and the gist of the case if it's on the website, otherwise null

Return only JSON, without markdown.
"""
        }]
    )
    
    try:
        return json.loads("".join(b.text for b in message.content if b.type == "text"))
    except json.JSONDecodeError:
        return {
            "value_prop": "N/A",
            "differentiator": "N/A",
            "case_study": None
        }


def export_to_csv(leads: list[dict], filename: str = "leads.csv") -> None:
    """Step 4: Export to CSV"""
    if not leads:
        print("No leads to export")
        return
    
    fieldnames = ["company_name", "location", "website", "phone",
                  "value_prop", "differentiator", "case_study"]
    
    with open(filename, "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=fieldnames)
        writer.writeheader()
        for lead in leads:
            # Make sure every field is present (protects against shifted columns)
            row = {field: lead.get(field, "") for field in fieldnames}
            writer.writerow(row)
    
    print(f"Exported {len(leads)} leads to {filename}")


def run_pipeline(niche: str, location: str, count: int, output_file: str) -> None:
    """Run the full pipeline"""
    print(f"Starting: {niche} in {location}, {count} results")
    
    # Step 1: Scraping
    raw_leads = scrape_google_maps(niche, location, count)
    
    # Step 2: Filtering
    filtered_leads = filter_leads(raw_leads)
    
    # Step 3: Enrichment
    enriched_leads = []
    for i, lead in enumerate(filtered_leads):
        print(f"  Enriching {i+1}/{len(filtered_leads)}: {lead['company_name']}")
        website_text = scrape_website(lead["website"])
        enrichment = enrich_with_claude(lead["company_name"], website_text)
        enriched_leads.append({**lead, **enrichment})
    
    # Step 4: Export
    export_to_csv(enriched_leads, output_file)
    print("Done!")


if __name__ == "__main__":
    run_pipeline(
        niche="roofing companies",
        location="North Carolina",
        count=30,
        output_file="leads.csv"
    )

Real output: what ends up in the CSV

After a run, the system produces a file with data like this:

company_name location website phone value_prop differentiator case_study
Example Roofing Co Charlotte, NC example.com +1-555-0100 Full-service roofing with a family approach and financing Family-owned; years in business listed on the website Client: full roof renovation finished with no delays
... ... ... ... ... ... ...

The row above is made up, just as an example. This CSV goes into Google Sheets through File → Import (if everything lands in one column, Data → Split text to columns will fix it).


Packaging it as a skill + slash command

🎨 Picture this: Packaging it as a skill is like turning your recipe into a dish on a restaurant menu. You wrote the recipe yourself. The menu means any server can take the order and the kitchen knows what to do without you in the room.

Once the system works, you ask Claude Code to turn it into a skill:

Type this into the chat
Now that I've gone through this whole pipeline, create a skill
called "lead-generation-system" so I can call it
as the slash command /lead-generation-system.

Claude creates this structure:

Code
.claude/skills/lead-generation-system/
├── SKILL.md          ← description + instructions for Claude
├── scripts/
│   └── lead_generation.py
└── references/
    ├── requirements.txt
    └── sample_output.csv

The skill becomes a slash command by itself: now you can type /lead-generation-system and the system will ask for the niche, location and count. A separate file in .claude/commands/ is no longer required: commands and skills have been merged, and old command files keep working. To make the skill available in every project, put it in your personal ~/.claude/skills/ folder.


Ethical limits

OK Not OK
Public data from Google Maps Email addresses of private individuals
Public text on company websites Owners' personal data
Business phone numbers from public sources Aggressively bypassing CAPTCHAs
Analyzing public case studies Data from private sections of websites

The system collects only what businesses have publicly posted online themselves, which is usually acceptable for B2B outreach. Rules on personal data and email marketing depend on the country (for example, CAN-SPAM in the US, CASL in Canada, GDPR in Europe): check the law where you and your recipients are before sending anything. This is not legal advice.


Practice

Assignment: Lead generation for a specific niche

  1. Sign up at apify.com and get an API token (the free plan includes a small credit; current terms are on the Apify website)
  2. Install the dependencies: pip install apify-client anthropic requests beautifulsoup4
  3. Ask Claude Code to build a simplified version of the pipeline for a niche you choose:
    • "digital marketing agencies in Miami"
    • "yoga studios in Austin"
    • IT companies in your city (in a niche you understand)
  4. Run it with count=10 as a test and make sure the CSV is generated correctly
  5. Open the CSV in Google Sheets and check that all the columns are in the right place
  6. Ask Claude Code to package the system as a skill
  7. Bonus: add a scoring column (0-100) where Claude rates the lead's quality based on whether there's a case_study and how detailed the value_prop is

Goal: go from zero to a working lead generation system with real data.


Tools and resources

  • Apify: apify.com, a marketplace of scrapers; Google Maps scraper ID: nwua9Gu5YrADL7ZDj
  • apify-client: pip install apify-client, the Python SDK for Apify
  • beautifulsoup4: pip install beautifulsoup4, parses website HTML
  • requests: pip install requests, HTTP requests to the leads' websites
  • Google Sheets: Data → Split text to columns, to open the CSV
  • Hunter.io: hunter.io, finds emails by domain (there's a free plan with a monthly credit limit; see the website), adds contact details to the pipeline
  • Apollo.io: apollo.io, a full B2B contact database with email verification, an alternative to Apify for leads
  • LinkedIn Sales Navigator: business.linkedin.com/sales-solutions, targeted search for decision-makers (see the website for trial terms)
  • Instantly.ai: instantly.ai, automated email sending with domain warm-up (for when you scale your outreach)

Common mistakes

🎨 Picture this: Not filtering leads before enrichment is like refining ore together with the gravel. More expensive, slower, and half the effort goes into the trash. Filter first, then enrich only what's valuable.

  • Not filtering leads before enrichment. Enriching each lead costs money (an API call to Claude). Without a filter, you spend your budget on companies with no website or phone that you can't contact anyway. Filter BEFORE enrichment.
  • Not checking the CSV by hand. Always eyeball the first 10 rows in Google Sheets. Shifted columns, empty fields, duplicates: all of this is easy to catch early, but at scale it ruins the whole database.
  • Too large a count at the start. Start with count=10 as a test. If you go straight to 500, you'll spend money on Apify and the Claude API, spend a lot of time on enrichment, and then discover a bug in step 3. Check the cost and data quality on a small run first.
  • Not connecting leads to outreach. A CSV without action is just a file. This pipeline feeds cold outreach (→ Cold outreach with Claude) and your CRM (→ the Google Sheets table from the same lesson).

How this connects to other lessons


Key takeaways

A 4-step pipeline: Apify (scrape) → filter (website + phone) → Claude enrichment → CSV. Each step is independent and tested separately.

Build and test the system first, then ask Claude to package it as a skill. A skill made from a working system beats a skill made "out of thin air."

Protect against shifted CSV columns: set fieldnames explicitly and use lead.get(field, ""). This prevents the phone number from landing in the location column.


Next lesson

→ Pricing: value-based pricing

The mark stays in this browser only and is never sent anywhere. My progress