SISuperintelligenceDocs

Search docs

Search every page of the documentation.

Self-hosted agents

Self-hosted agents

A small program on your own machine that runs HTTP requests, browser pages and shell commands for your organization, within the permissions you approve.

Some work has to come from a machine you control: sites that block cloud addresses, services on your network, or tools installed on your computer. The agent (si-agent) runs on macOS, Linux or Windows, connects to your organization, and runs jobs the organization queues.

How it works

si-agent login   ──▶  you approve the device at id.dev.gov.vin: organization, capabilities, network
si-agent run     ──▶  keeps a connection open that announces new jobs, and polls every 60 s
                       claims a job it's allowed to run, runs it on this machine
                       sends a heartbeat every 20 s (cancels arrive with it)
                       uploads the results, reports done, failed or canceled
platform          ──▶  delivers the results to the job's callback URL, if it has one
  • The agent only connects out to the platform; nothing connects in.
  • It runs one job at a time.
  • It runs as the user who installed it, as a background service that starts at login, and installs new signed releases by itself.

Capabilities

When you approve a device you choose what it may run:

CapabilityJob typeDoes
HTTP requestshttp.batchFetches URLs.
Browser pagesbrowser.batchLoads pages in the Chrome, Chromium or Edge installed on the machine, and extracts data from them.
Shell commandsexecRuns commands with the user's privileges. On an allowlist, only where the agent can sandbox them (macOS, Linux).

Network access

Each device has a network mode:

  • Specific hosts (allowlist): HTTP and browser jobs can reach only the listed hosts. When a job needs another host, the agent asks for it, and the request shows up in Cloud → Agents to allow or deny.
  • Full access: jobs can reach any host.

The agent enforces the allowlist itself: on each HTTP request and the URL it ends up at after redirects, and on every page a browser job navigates to. A page's own scripts, images and requests aren't checked. Shell commands on an allowlist run in an operating system sandbox that can only reach the allowed hosts, through a proxy the agent runs (details).

HTTP and browser jobs obey robots.txt and identify themselves.

Pages