SISuperintelligenceDocs

Search docs

Search every page of the documentation.

Self-hosted agents

Browser jobs

How browser.batch loads pages: which browser, how it identifies itself, robots.txt, pacing and captures.

See Job types for the input and output.

The browser

  • The agent uses the Chrome, Chromium or Edge already installed on the machine; it never downloads one. Set SI_AGENT_CHROME to the browser's executable to choose one yourself.
  • Each job gets a fresh browser with a temporary profile, so no cookies or sign-ins carry over between jobs. The browser is closed when the job ends.
  • Pages load headless in a 1280 × 900 window, with downloads blocked and the popup blocker on.

Identification

Browser jobs don't disguise themselves:

  • The user agent is the browser's own followed by si-agent/<version>.
  • The browser's automation flag stays on (navigator.webdriver is true).
  • There are no stealth plugins, fingerprint changes or CAPTCHA solving. A site that challenges automated browsers gets its challenge page back.

robots.txt

Before a site's first page in a job, the agent reads its /robots.txt (just the file: no scripts, and a redirect to a host that isn't allowed doesn't count):

robots.txtResult
200 with rulesPages must be allowed by the rules.
404 or 410No restrictions.
Anything else: 401, 403, 5xx, an HTML page, a timeout, no connectionThe whole site is skipped for this job.

How rules apply:

  • Rules in the * group and in a group for si-agent both apply: a path must be allowed by each.
  • Within a group, the longest matching rule wins, and a tie between allow and disallow allows.
  • * matches any characters and a trailing $ anchors the end, as in Google's matcher. Paths include the query string.
  • Only the first 500 KB of the file is read.

crawl-delay is honored: the longest one from the groups that apply sets the pace. A site asking for more than 60 seconds is skipped.

Disallowed pages return robots_disallowed, and responses from disallowed paths aren't captured.

The allowlist

On a device with an allowlist, every page navigation is checked: the page itself, each redirect, and navigations the page starts. A navigation to a host that isn't allowed is blocked, the page returns network_denied, and the agent files a network request for the host.

The scripts, images and requests a page makes while loading aren't checked, but captures are kept only from allowed hosts. Service workers are bypassed, so they can't answer navigations.

Pacing

Pages load one at a time. Before each page the agent waits the larger of:

  • delayMs (default 8 seconds, kept between 5 and 60 seconds), and
  • the site's crawl-delay.

The wait counts from the previous page, and from the last page of the same site in an earlier job on this device, so back-to-back jobs for one site keep the same pace.

Loading a page

Each page gets 45 seconds, 5 of them reserved for extraction:

  1. Open the URL and wait for the document.
  2. Wait for the page to finish loading (up to 15 seconds).
  3. Wait for waitForSelector, if given, until the page's time runs out; a missing selector ends the page with timeout.
  4. Scroll down a screen at a time, if scroll is set, until the bottom of the page (up to 10 seconds).
  5. Wait waitMs more, if given (up to 20 seconds).
  6. Extract the requested data and collect captures.

Captures

With capture, the agent keeps responses the page receives whose URL matches urlPattern, a JavaScript regular expression, for example /api/products\?:

  • Only from allowed hosts and paths robots.txt allows.
  • Not redirects or OPTIONS requests.
  • Up to max responses (default 10, at most 50), each body cut at 5 MB.

Captures are collected until extraction starts.

Size of the results

A job's results can reach about 50 MB. After that, the remaining pages aren't loaded and return output_limit; queue them in another job.

Try a page

browser-test loads one page through the same code, with full network access and every extraction, and prints what it got. robots.txt still applies.

si-agent browser-test https://www.example.com/products --selector "main" --capture "api/products" --wait 2000 --scroll

It prints the status, the final URL, the size of each extraction and the captured responses.