Skip to main content
An environment is the saved definition of the sandbox a session runs in: its compute size, workspace disk, setup script, and network access. You create it once, prove its setup script works with a setup run, and then name its environment_id when you start sessions. Creating an environment starts no compute. Compute starts only when you run a setup test or start a session. Each top-level session gets its own fresh sandbox, so two sessions on the same environment never share files. Subagents and teammates that a session starts share that session’s sandbox. For every field, default, and limit, see Environment reference.
Creating, changing, testing, or deleting environments needs the organization developer or admin role. The organization user role can view environments and their setup runs. See Organizations and roles and API keys. Setup runs start real compute, the same kind a session uses. New environments start with no internet access, so decide first which hosts the sandbox must reach, such as package registries or github.com.

Create an environment

This example creates an Agent runner environment with 2 vCPU, 4 GiB of memory, a 20 GiB workspace, and a setup script that installs two pinned Python packages. Its network policy allows only the two Python package hosts the script needs. Creating the environment saves only the configuration, so the example then starts a setup run explicitly.
  1. In the sidebar, click Environments, then click Create environment.
  2. Enter a Name. The form uses Agent runner automatically. To start from a template instead, see Start from a template.
  3. Under Permissions, set Internet access to Enabled. The console offers only Enabled or No access. To allow only specific hosts, use the API as in the cURL tab, or edit the policy later as shown in Control network access.
  4. Expand Compute, set Compute kind to CPU, and click Standard (2 vCPU, 4 GiB).
  5. Expand Storage and lifecycle and set Workspace disk (GiB) to 20.
  6. Paste the script into Setup script.
  7. Click Save and test setup. The Setup test panel streams the run’s log.
The create call returns 200 OK with the saved environment. It starts no compute. Because the environment has a setup script but no completed setup run yet, setup_verification.status is never. The setup-run call returns 202 Accepted with a separate run whose initial status is queued. Keep environment_id to start sessions and setup_run_id to follow the setup run.
  • Saving and testing are separate. createEnvironment stores the configuration. createEnvironmentSetupRun provisions compute and requires its own Idempotency-Key.
  • Create retries can be idempotent. Idempotency-Key is optional on createEnvironment. Reuse the same key and exact request on an ambiguous retry to replay the first response instead of creating another environment.
  • Network access is closed by default. Omitting network_policy stores {"version":"v1","rules":[]}, which blocks every host. Send {} only when you deliberately want unrestricted outbound access.
  • Privileged Docker is off by default. privileged defaults to false. To set it to true, send network_policy: {} for unrestricted internet access.
  • Every environment starts from the Agent runner image. Add tools and data with the setup script. See What the sandbox starts with.
A runtime (the provider field) is the kind of sandbox the environment runs on. The console’s Create environment form uses the catalog’s default runtime automatically, and the runtime cannot be changed after creation. API clients should list the available runtimes rather than hard-coding one. Today the list contains runs: managed Linux compute with a persistent workspace, network policy, and setup verification.
The response also carries a description for each runtime. The list is the same for every organization and supports If-None-Match for caching.

Start from a template

Templates are reviewed starting points for common work. A template fills the create form with its compute, network policy, and setup script, if it has one. Review the settings, change what you need, then create the environment. Base sandbox and Browser desktop have no setup script, so the environment is ready as soon as you create it. A template with a setup script, like any environment with one, needs a passing setup test before sessions can start.
  1. In the sidebar, click Environments, then click Create environment.
  2. Under Choose a starting point, click a template. The form fills in.
  3. Review the settings. A template that lists hosts keeps its own network policy, limited to the package and source hosts it needs.
  4. If the template has a setup script, click Save and test setup. Otherwise, click Create environment. The Setup column reads Ready.
listEnvironmentTemplates returns each template under items with its id, revision, name, summary, and a definition you can send to createEnvironment as the request body. Creating from a template works like any other create: if the definition has a setup script, start a setup run and wait for it to pass before you start sessions. Without one, sessions can start right away.

Write the setup script

The setup script runs on fresh compute after the sandbox starts and before the agent’s first turn. It runs as the sandbox user in a bash login shell with set -eo pipefail. In a setup run, it gets the environment’s variables and network policy, but no session credentials such as vault grants.
  • One command per line. The first failing line ends the run, and the result names that line.
  • Don’t wrap lines in bash -lc. The script already runs under bash, and the wrapper’s quoting is a common failure.
  • Pin versions. The run proves the exact script, so unpinned installs can drift from what was proven.
  • Put Python packages in /workspace/.venv. The sandbox selects that virtual environment for the agent’s commands automatically. Install through its interpreter, for example uv pip install --python /workspace/.venv/bin/python.
  • Don’t rely on export, source, aliases, or shell functions. The agent’s commands run in a new shell, so none of these carry over.
  • Use sudo for apt. The script doesn’t run as root.
  • Put files the agent needs under /workspace. They stay with the session after its sandbox stops or is released and can be used when its workspace reopens.
The script can be at most 64 KiB. setup.timeout_seconds bounds the whole script; it defaults to 600 and accepts 10 through 3600. A script that runs past the timeout fails with exit code 124.The service saves the script with Windows line endings converted and trailing whitespace removed. setup_warnings flags risky lines such as curl | sh, unpinned installs, and bash -lc wrappers. Warnings are advice; they never block a save or a run.

Verify the setup script

An environment without a setup script is usable as soon as you create it. An environment with a setup script is usable only after a setup run passes for its current configuration. A setup run starts fresh compute from the environment, runs the setup script exactly as a session would, records every line of output, checks the compute, saves a reusable image, and releases the compute. The result becomes the environment’s setup_verification.
A session can start on the environment when either of these is true:
  • setup_verification.status is not_applicable, because there is no setup script. In the console, the Setup column reads Ready.
  • setup_verification.status is verified, stale is false, and the service kept a reusable image of the prepared sandbox. In the console, the Setup column reads Verified.
Otherwise startSession returns 422 environment_not_verified, and details.next_action names the call that fixes it.When a setup run passes, the service saves the prepared sandbox as a reusable image. New sessions start from that image, so they don’t repeat the setup work. If the runner changes, the platform automatically runs setup on dedicated compute, captures a new image, and keeps new sessions in provisioning until that image is ready. The session page shows the setup-run log while it waits. A failed refresh fails the waiting session and links its setup log from the session failure. The environment retains its prior verified verdict, but sessions needing the new runner image remain blocked until setup is tested again.

Start a setup run

createEnvironmentSetupRun requires an Idempotency-Key. Use one key for each logical setup run and reuse it with the same request if the response is lost. Only one setup run can be in flight per environment, and your organization can have 4 manual setup runs in flight at once.
  1. In Environments, open the environment.
  2. In the Setup test panel, click Test setup.
The call returns 202 Accepted with the queued run. Some fields are omitted here.
If a run is already in flight, the call returns 409 setup_run_in_progress with that run’s id in details.setup_run_id. An environment with no setup script returns 400 invalid_request, because there is nothing to verify.

Wait for the result

Poll getEnvironmentSetupRun until status is succeeded, failed, or cancelled. Pass wait_seconds to have the server hold the request until the run finishes or the wait ends. The server caps each wait at 5 seconds, so loop until the status is terminal.
  1. Watch the Setup test panel. It shows each phase and streams the log.
  2. The status reads Setup verified. on success, or shows the hint and failing line on failure.
Some fields are omitted here. next_action names the call to make next, such as polling again or reading the log. What success means for a setup run. status is succeeded, and the environment now reads setup_verification.status: "verified" with stale: false for the configuration the run tested. Sessions can start. If the script passed but the image step failed, the console shows Setup passed, but its reusable image failed. and sessions still return environment_not_verified; run setup again. If the run failed, see Fix a failed setup run. A setup run moves through these statuses:
phase tells you where the run is or where it ended: provision, gpu_check, setup, profile, commit (saving the reusable image), or cleanup. Only a failure in setup points at your script.
  • failed records a failed verdict. Sessions can’t start until a later run passes.
  • cancelled records nothing. If the environment was verified and unchanged when the run started, it stays verified. If it read running, it goes back to never.
  • Re-verifying doesn’t interrupt sessions. While a new run is in flight on a verified, unchanged environment, it stays verified and sessions keep starting. The console shows Re-testing. In every other case, starting a run sets the status to running, replacing an earlier failed or stale verdict.
The log interleaves four streams: stdout, stderr, marker (the script line about to run), and system (notes about phases outside your script). Secrets are redacted before the log is stored.In the console, open the environment. The Setup test panel streams the log of the selected run. Under Recent runs, pick an earlier run to read its log.With the API, read from the start with after=0. Each page returns up to limit lines (default 500, maximum 2000) and a next_after cursor. To tail a live run, pass next_after as after on the next call with wait_seconds=0, and wait about two seconds between calls. Stop when status is terminal and a page comes back empty.
The service keeps up to 4 MiB of log per run. Past that, truncated is true; the beginning of the log is kept, and the last 2 KiB of stderr stays in the run’s stderr_tail.
listEnvironmentSetupRuns returns the environment’s recent runs, newest first. It includes manual runs (kind: "manual", which you start), automatic runner refreshes (kind: "runner_refresh", shared by sessions that need an updated setup image), and session runs (kind: "session", recorded when a session ran setup on its own compute, with its session_id). The newest manual or runner refresh run is the one setup_verification describes. limit accepts 1 through 50 and defaults to 50. In the console, open the environment and read the Recent runs list in the Setup test panel.
setup_runs is an empty array when no run has been requested.
A member allowed to update the environment can cancel a manual run or automatic runner refresh that is queued, provisioning, or running. The run stops, its compute is released, and active_setup_run_id is cleared. A cancelled run records no verdict. In the console, open the environment while a run is in flight and click Cancel run in the Setup test panel. The status reads Cancelled.
A run that already finished returns 409 setup_run_finished, and nothing changes. A session run returns 400 invalid_request; stop the session instead.

Fix a failed setup run

On a failed run, read hint first. Then read failed_line, failed_command, exit_code, and stderr_tail. Fetch the full log only if those don’t explain the failure. The same summary appears in the environment’s setup_verification.last_run. When phase is anything other than setup, the failure happened outside your script; retry once before you change anything. failure_code classifies the failure with the same codes session failures use, such as environment_setup_failed, sandbox_capacity_unavailable, or sandbox_provision_timeout.

Control network access

network_policy controls which hosts the sandbox can reach. It applies to the setup script, the agent’s commands, package downloads, and websites opened by the sandbox browser. Outbound traffic that no rule allows is blocked. Rules are checked in order, and the first match wins. For the rule grammar, see Network policy.
Access is granted per host, not per action. Once a host is allowed, sandbox code can send anything to it, including uploads such as git push. Allow only the hosts the work needs, especially when agents process untrusted content.
Some traffic doesn’t leave from the sandbox, so the policy doesn’t govern it: model calls, calls to the agent’s MCP servers, and the web_search and web_fetch tools. Control those in the agent’s configuration. See Tools.
  • Automated sessions in a restricted environment don’t get web_fetch. A session started by an automation or from Slack, in an environment with limited or no internet access, has no web_fetch tool, so outside text can’t use it to send data out. Its sandbox commands still reach the hosts the policy allows.
  • Built-in integration apps reach their own hosts. The tools of apps from the Integrations catalog, such as Jira and Confluence, run inside the sandbox. Recursion allows the hosts they need, so they work with Enabled internet access or a policy limited to specific hosts. Only No access blocks them.
Privileged Docker requires the Enabled mode. A nonempty policy, including No access and custom host rules, cannot be saved together with privileged: true on create or when access settings change.

Allow specific hosts

A policy change must prove you saw the current access settings. Read the environment, then send its current network_policy and privileged values in expected_access along with the new policy. This example allows the hosts that git and the gh CLI use for GitHub. Attaching a GitHub credential in a vault doesn’t open these hosts; the environment’s policy must allow them.
The console can’t edit individual host rules. Under Internet access, choose Enabled for unrestricted access or No access to block everything, then click Save environment. Choosing either replaces a saved custom policy. Use the API for host rules.
The PATCH returns the updated environment. Because the policy changed, the earlier verdict is stale. Start a setup run and wait for it to pass before starting sessions.
Some fields are omitted here.
  • network_policy in a PATCH replaces the whole policy. Include every rule you want to keep.
  • expected_access is required when a PATCH changes a saved policy other than {} (including the default no-internet policy), or turns privileged on. Without it the API returns 428 precondition_required. If either value changed since your read, it returns 412 precondition_failed; read the environment again and retry.
  • Closing access (changing {} to a restricted policy, or turning privileged off) doesn’t need expected_access.
  • Omitting network_policy or privileged in a PATCH keeps the saved value.
  • A PATCH that changes access settings cannot combine privileged: true with a nonempty policy; the API returns 400 invalid_request. An existing environment with both settings can still be viewed and have other settings edited. To start a new session, set network_policy: {} or turn off privileged Docker.

Update an environment

An update is partial: omitted fields keep their saved values. Changes apply only to sandboxes started afterward; a session already running keeps the configuration it started with. This example raises compute to 4 vCPU and 8 GiB and raises the idle stop to one hour.
  1. In Environments, click the environment name, or choose Edit from its action menu.
  2. Expand Compute and click Heavy (4 vCPU, 8 GiB).
  3. Expand Storage and lifecycle and set Idle stop after (seconds) to 3600.
  4. Click Save and test setup, or Save environment to save without testing.
The PATCH response is the updated environment. Some environment fields are omitted here.
Changing compute makes the earlier verdict stale. After a verification-affecting change, start a setup run explicitly and wait for it to pass before starting new sessions. Changing only the name, description, metadata, idle stop, retention, or computer use keeps the verdict. The full list is in Setup verification.
A few fields replace the saved value as a whole:
  • setup replaces the whole setup object. Send both script and timeout_seconds to keep both. {"script": ""} removes the script.
  • resources replaces the whole compute size.
  • env_vars and mounts replace the whole map or list. Send {} or [] to clear them.
A concurrent edit to the same environment’s access or compute can return 409 revision_conflict. Read the environment again and retry.

List and get environments

listEnvironments returns every environment in your organization, newest first, with the full configuration. getEnvironment returns one environment.
  1. In the sidebar, click Environments. The table shows each environment’s Name, ID, Resources, and Setup status.
  2. Search by name or ID.
  3. Click a name, or choose Edit from its action menu, to open the full configuration.
The list is not paginated, and deleted environments are left out. An unknown, deleted, or other-organization id returns 404 not_found from getEnvironment. environments is null, not an empty array, when the organization has no environments. Some fields are omitted here.

Delete an environment

Deleting removes the environment from lists and reads, and new sessions can’t use it. Sessions that are running keep their sandbox, and past session history stays readable. There is no restore, so create a new environment if you need it back.
  1. In Environments, open the environment’s action menu.
  2. Click Delete.
  3. In the Delete environment dialog, click Delete.

What can go wrong

The most common problems:
  • 422 environment_not_verified on startSession. The setup script hasn’t passed a current setup run. Follow details.next_action: start a setup run, or wait for the one in flight.
  • A setup run fails with egress_blocked. The network policy blocks a host the script needs. Allow the host, then run setup again.
  • The agent’s git clone of a public repository fails. The policy doesn’t allow the GitHub hosts. Allow the hosts shown in Allow specific hosts.
  • 428 precondition_required or 412 precondition_failed on a PATCH. Access changes need the current network_policy and privileged in expected_access. Read the environment and resend.
For the full error catalog and retry guidance, see Errors.

Limits

An environment name can be up to 256 characters. Setup script, setup-run, log, and network policy limits are in Limits. For compute, disk, and lifecycle limits, see Environment reference.

Next steps

Environment reference

Every field, default, compute size, sandbox path, and network rule.

Start a session

Combine a usable environment with an agent and start work.

Authenticate with vaults

Give agents credentials without putting secret values in the environment.

Deliverables and artifacts

Where agents save the files you asked for.