Library · Autonomous and multi-agent systems

AI sandboxes: E2B and Modal for running agent code safely

Engineer60 minUpdated: October 2026
72 of 105 in the library

Time: about 25 min reading + 35 min practice


The gist

When an agent writes and runs code, who's responsible for whatever that code does? If the agent runs rm -rf / or accidentally deletes a database, the damage lands on your machine, on your client's systems, in production.

Sandboxes solve this problem at the root: the agent gets an isolated environment, a virtual container, where it can run any code, read and write files and make network requests. But all of that happens inside a "cage." The world outside stays untouched.

🎨 Picture this: A sandbox is like the glove box in your car where you keep everything that might leak, spill or explode. The car itself stays clean.

In this lesson we'll look at the two main tools: E2B, sandboxes built specifically for AI agents, and Modal, a serverless platform for ML tasks and heavy computing.


Key concepts

  • Sandbox: an isolated environment for running code with limited access to outside resources
  • Containerization: a technology (Docker) that isolates processes at the operating system level
  • Stateful sandbox: a sandbox that keeps its state between calls (files, installed packages, variables)
  • Cold start: the time it takes to start a sandbox from scratch; for E2B it's on the order of a second (exact numbers are in the E2B documentation)
  • Filesystem access: the agent can read and write files inside the sandbox, just like on a regular computer
  • Network isolation: a sandbox may or may not have internet access (depends on settings)
  • Pay-per-use: both tools bill for actual usage time, not for a fixed server
  • GPU sandbox: Modal gives you GPUs for ML tasks without renting expensive servers

Theory

Why agents need sandboxes

Imagine you hired a programmer who works remotely. You trust them, but their machine has your data, your API keys, your production database. Even an honest mistake can be expensive.

With AI agents the situation is sharper: the agent doesn't just write code, it runs it. And if the agent misunderstood the task, wrote a destructive script, or fell for a prompt injection, then without isolation the consequences instantly spill beyond what you intended.

Three main scenarios where you need isolation:

  1. Running untrusted code. A coding agent that writes and tests code for a user's task. The code may be logically correct but dangerous in its side effects.

  2. Analyzing data without risk. An analyst agent that gets a CSV, an Excel file or a database to explore. It's better for it to work with a copy inside a sandbox, not with the original.

  3. Testing before production. A deployment agent checks that the build works and runs the tests, and only then gives the go-ahead for the real deploy.

🎨 Picture this: A sandbox for an agent is like a chemist's lab. You can mix anything there, explosions are acceptable, and that's exactly why there's a fume hood, an apron and safety goggles. You don't carry the chemicals into your living room.

E2B: sandboxes built specifically for AI agents

E2B (e2b.dev) is a company that did one thing and did it well: micro virtual machines where agents run code.

What defines E2B:

  • Fast start, on the order of a second. A regular Docker container usually takes noticeably longer to start. E2B uses lightweight virtual machines based on Firecracker (an AWS technology), which spin up an isolated environment very quickly.

  • Python and JavaScript SDKs. You control the sandbox through code: create it, run a command, get the result, close it.

  • Persistent filesystem. Files the agent creates in the sandbox are kept for the whole session. You can download the result afterward.

  • Network access. By default the sandbox can reach the internet, so the agent can download packages and make API requests.

  • Custom templates. You can create your own template with preinstalled dependencies so you don't have to install them every time.

Example: a Python coding agent with E2B:

python
import anthropic
from e2b_code_interpreter import Sandbox

client = anthropic.Anthropic()

# Create a sandbox
sandbox = Sandbox.create()

# The prompt for the agent
user_task = """
Write a Python script that:
1. Generates 100 random numbers from 1 to 1000
2. Finds the mean, median and standard deviation
3. Plots a histogram and saves it to histogram.png
Run the script and show the results.
"""

# The agent generates the code (current models: the "What's current" page)
response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=2048,
    messages=[{"role": "user", "content": user_task}]
)


def extract_code(text: str) -> str:
    """The model often wraps code in ```python ... ```: pull out just the code."""
    if "```" not in text:
        return text
    block = text.split("```")[1]
    return block.removeprefix("python").strip()


# Pull the code out of the agent's response
agent_code = extract_code("".join(b.text for b in response.content if b.type == "text"))

# Run the agent's code in the isolated sandbox
execution = sandbox.run_code(agent_code)

# Get the results
print("Stdout:", execution.logs.stdout)
print("Stderr:", execution.logs.stderr)

# Download the result file
if execution.results:
    for result in execution.results:
        if hasattr(result, 'png'):
            with open("histogram.png", "wb") as f:
                f.write(result.png)
            print("Histogram saved to histogram.png")

# Close the sandbox (frees all resources)
sandbox.kill()

Ways to use E2B:

  • AI coding assistant: the user asks for a script, the agent writes the code, tests it in the sandbox, fixes errors and returns a working result
  • Data analysis agent: you upload a data file, the agent analyzes it, builds charts and returns insights
  • Web scraping agent: the agent writes and tests a scraper in a safe environment, with no risk of getting your production IP blocked
  • Code review with execution: the agent doesn't just read the code, it runs the tests and checks edge cases

Modal (modal.com) is a different class of tool. It's not just a sandbox for code, it's a full serverless platform with a focus on ML tasks.

What Modal can do that E2B can't:

  • GPUs on demand. You can run a function on L4, A10, A100, H100 and other GPUs and pay only for GPU time. No dedicated server needed.

  • Scheduled jobs. Run functions on a schedule (cron), like GitHub Actions but with access to GPUs and a Python environment.

  • Web endpoints. Turn Python functions into API endpoints with a single command.

  • Volume storage. Persistent storage for models, datasets and results.

  • Parallel jobs. Run thousands of jobs in parallel without managing infrastructure.

🎨 Picture this: If E2B is the glove box for an agent's risky experiments, Modal is an entire research institute with a supercomputer, labs and a schedule. You use only what you need and pay for the time you actually use.

Example: running an ML model on a GPU with Modal:

python
import modal

app = modal.App("ai-inference")

# An image with the dependencies
image = modal.Image.debian_slim().pip_install(
    "torch", "transformers", "pillow"
)

@app.function(
    image=image,
    gpu="A10",           # GPU type: see the Modal documentation for allowed values
    timeout=300
)
def run_inference(prompt: str) -> str:
    from transformers import pipeline
    
    generator = pipeline("text-generation", model="gpt2")
    result = generator(prompt, max_length=200)
    return result[0]["generated_text"]

# Run from a local script (Modal spins up the GPU instance itself)
if __name__ == "__main__":
    with modal.enable_output():
        with app.run():
            result = run_inference.remote("Write a short story:")
            print(result)

E2B vs Modal: when to use which

Criterion E2B Modal
Main job An agent runs user code ML tasks, heavy compute
Startup speed on the order of a second seconds, depends on the image and GPU (cold start)
GPU No Yes (L4, A10, A100, H100 and others)
Stateful filesystem Yes, within a session Yes, through Volume
Scheduled jobs No Yes
Web endpoints No (API only) Yes
Target user AI agent developer ML engineer, data scientist
Free start Yes (one-time credits) Yes (monthly credits)
Billing per second for vCPU and memory per second for cores, memory and GPU

Bottom line: E2B is for agents that write and run arbitrary code. Modal is for agents that need serious computing power (inference, fine-tuning, batch processing).

Other alternatives

Daytona: an open-source, self-hosted alternative to E2B. If privacy matters (you're running clients' code), you can host it on your own server. It supports the same operations: creating an environment, running code, a filesystem.

Docker directly: if you want full control and are ready to set it up yourself. More flexibility, but more work. No built-in integration with AI agents.

GitHub Codespaces / Gitpod (now Ona): development environments, not agent environments. Too heavy and too expensive for per-request use.

Security: what an agent CAN and CANNOT do in a sandbox

Can:

  • Read and write files inside the sandbox
  • Run any shell commands (bash, python, node...)
  • Install packages with pip, npm, apt
  • Make outbound HTTP requests (if the network is allowed)
  • Use CPU and RAM within the limits

Cannot (isolated):

  • Access files on your machine
  • Read the host's environment variables
  • Change anything outside the container
  • Use more CPU/RAM than it's been given

An important note: E2B allows network access by default. If an agent in the sandbox receives a malicious prompt and wants to exfiltrate data, it can technically make an HTTP request to an outside server. For high-security scenarios, E2B supports network isolation.

🎨 Picture this: A sandbox is a prison with an exercise yard. The prisoner can't get past the walls, but with a smuggled note they can ask an accomplice outside to do something. Full security also requires monitoring outgoing requests.

Cost

Both services bill by the second for actual use, so you don't pay for idle time, and both let you start for free. E2B has a free Hobby plan with one-time credits when you sign up and limits on concurrent sandboxes and session length, and a Pro plan with a monthly fee plus usage, longer sessions and more concurrent sandboxes. Modal has a Starter plan with no monthly fee that includes credits every month; GPU time costs noticeably more than CPU time, and sandboxes cost more than regular functions.

Rates change, so we don't give figures here: before you plan a budget, check the current ones on the official pages, e2b.dev/pricing and modal.com/pricing. To estimate a task, multiply its duration in seconds by the resources it uses (vCPU, memory, GPU) and by the published rate. A short agent task on E2B's default configuration usually costs a fraction of a cent, but run the numbers with today's prices.

For prices and versions we track, see What's current.


Practice

Assignment: Build a data analysis agent that takes a CSV file, runs the analysis in an E2B sandbox and returns a report.

Step 1: Install and set up

bash
# Install the SDKs
pip install e2b-code-interpreter anthropic

# Get API keys
# E2B: https://e2b.dev/dashboard → API Keys
# Add to .env:
# E2B_API_KEY=e2b_...
# ANTHROPIC_API_KEY=sk-ant-...

Step 2: Create a basic sandbox

python
from e2b_code_interpreter import Sandbox
import os

# Create a sandbox
sbx = Sandbox.create(api_key=os.environ["E2B_API_KEY"])

# Check that it works
result = sbx.run_code("print('Sandbox is running!')")
print(result.logs.stdout)  # ['Sandbox is running!\n']

# Install dependencies (once per session)
sbx.run_code("import subprocess; subprocess.run(['pip', 'install', 'pandas', 'matplotlib', 'seaborn'], capture_output=True)")

print("Sandbox ready")
sbx.kill()

Step 3: Upload data to the sandbox

python
from e2b_code_interpreter import Sandbox
import anthropic, os

def analyze_csv(csv_path: str) -> dict:
    sbx = Sandbox.create(api_key=os.environ["E2B_API_KEY"])
    
    # Upload the CSV to the sandbox
    with open(csv_path, "rb") as f:
        csv_content = f.read()
    
    sbx.files.write("/home/user/data.csv", csv_content)
    
    # Check it
    result = sbx.run_code("""
import pandas as pd
df = pd.read_csv('/home/user/data.csv')
print(f"Rows loaded: {len(df)}")
print(f"Columns: {list(df.columns)}")
print(df.dtypes)
    """)
    
    print("Data in the sandbox:")
    print(result.logs.stdout)
    
    return sbx  # return it for further work

Step 4: The agent analyzes the data

python
def run_ai_analysis(sbx, question: str) -> str:
    client = anthropic.Anthropic()
    
    # Get the data structure
    structure = sbx.run_code("""
import pandas as pd
df = pd.read_csv('/home/user/data.csv')
print(df.describe().to_string())
print("\\nFirst rows:")
print(df.head().to_string())
    """)
    
    data_info = "\n".join(structure.logs.stdout)
    
    # Ask the agent to write the analysis code
    response = client.messages.create(
        model="claude-sonnet-5-5",
        max_tokens=2048,
        system="""You are a data analyst. Write clean Python code for data analysis.
        The data is already loaded at /home/user/data.csv.
        Use pandas, matplotlib, seaborn.
        Save charts as /home/user/plot.png.
        At the end, print short conclusions with print().""",
        messages=[{
            "role": "user",
            "content": f"Data:\n{data_info}\n\nTask: {question}"
        }]
    )
    
    # Pull the code out of the response (the extract_code function from the first example)
    code = extract_code("".join(b.text for b in response.content if b.type == "text"))
    
    # Run the agent's code in the sandbox
    execution = sbx.run_code(code)
    
    analysis_output = "\n".join(execution.logs.stdout)
    errors = "\n".join(execution.logs.stderr)
    
    if errors:
        print(f"Warnings: {errors}")
    
    return analysis_output

Step 5: The full pipeline, with results

python
def full_analysis_pipeline(csv_path: str, question: str):
    print(f"Starting analysis: {question}")
    
    sbx = Sandbox.create(api_key=os.environ["E2B_API_KEY"])
    
    try:
        # Upload the data
        with open(csv_path, "rb") as f:
            sbx.files.write("/home/user/data.csv", f.read())
        
        # Install dependencies
        sbx.run_code("import subprocess; subprocess.run(['pip', 'install', 'pandas', 'matplotlib', 'seaborn', '-q'], capture_output=True)")
        
        # The agent analyzes
        result = run_ai_analysis(sbx, question)
        print("\n=== Analysis results ===")
        print(result)
        
        # Download the chart if there is one
        try:
            plot_data = sbx.files.read("/home/user/plot.png")
            with open("analysis_result.png", "wb") as f:
                f.write(plot_data)
            print("\nChart saved: analysis_result.png")
        except Exception:
            print("No chart was created")
            
    finally:
        sbx.kill()
        print("Sandbox closed")

# Run it
if __name__ == "__main__":
    full_analysis_pipeline(
        csv_path="sales_data.csv",
        question="Find the top 5 products by revenue and show sales trends by month"
    )

Tools and resources

Tool Link What it's for
E2B e2b.dev AI sandbox where agents run code
E2B Python SDK pip install e2b-code-interpreter SDK for Python
E2B JS SDK npm install @e2b/code-interpreter SDK for JavaScript/TypeScript
Modal modal.com Serverless for ML and heavy compute
Daytona daytona.io Open-source, self-hosted alternative to E2B
Docker docker.com Basic containerization when you need full control
Firecracker AWS open source The hypervisor under E2B's hood (to understand how it works)
E2B Docs docs.e2b.dev Official documentation with examples

Key takeaways

"An agent without a sandbox is a surgeon without gloves. Maybe it'll be fine. But why take the risk?"

"E2B solves one problem: giving an agent a safe place to run code. It's not overhead, it's the professional standard for production AI systems."

"Modal isn't a sandbox, it's a supercomputer on demand. If E2B is the glove box for experiments, Modal is a lab with GPUs. For ML agents, that difference matters."


Next lesson

→ Browser agents: Browserbase, Stagehand and web automation

Agents don't just write code; they can also drive a browser. In the next lesson we'll look at how to let an agent open websites, fill out forms, scrape data and automate web tasks with Browserbase and Stagehand.

The mark stays in this browser only and is never sent anywhere. My progress