ci is the HWS continuous-integration service: it plans workflow files from your repository into jobs, queues them on NATS JetStream, runs each job in a fresh heyvm microVM on one of your runner hosts, and serves a dashboard and machine API over the results.
What it is
ci is a single Rust process (ci/) with no agent to install on build machines. A machine becomes a runner by running heyvmd and joining a heyvm network; ci discovers those hosts through the heyvm control plane, opens an authenticated tunnel to each daemon, and drives builds on them.
| Piece | Role |
|---|---|
| Postgres | Authoritative state: runs, jobs, steps, logs, source descriptors, artifact records, repository registrations, VM ownership, deployment and release ledgers. |
| NATS JetStream | The job queue (<prefix>_JOBS, work-queue retention) and an event stream (<prefix>_EVENTS). Messages carry job IDs only. |
| heyvmd runners | Linux hosts in a heyvm network. Each job gets a new Firecracker or KVM microVM on one of them. |
| Native runners | Optional Intel macOS and Windows x86_64 hosts that poll ci over HTTPS and run jobs directly on the host. |
| heyosecret (optional) | Source of ${{ secrets.* }} and ${{ vars.* }}. See heyosecret. |
| artifacts store (optional) | Destination for ci/upload-artifact. See artifacts. |
| app-lb (optional) | Fronts the dashboard with sign-in, and stores workflow objects. See app-lb. |
The workflow format is GitHub Actions-shaped (jobs, needs, if, strategy.matrix, steps with run or uses, ${{ }} expressions), with two differences: a job's uses: names the heyvm network and host it runs on, and a job's vm: block describes the machine to boot.
You trigger runs with git submit, a Git subcommand that sends an exact revision to ci. There is no webhook integration with a Git host.
Architecture
From submit to result
- Submit.
git submitpicks a commit thatoriginalready has, computes a binary patch from it to the tree you want built, collects the workflow YAML from that tree, and POSTs agit-patchdescriptor to/api/submit. No repository archive is uploaded. - Plan. ci selects workflow files whose
on:includessubmit, applies each file's branch and path filters, expands matrices, resolves each job's network, and stores the expanded job plans and the source descriptor in Postgres in one transaction. Each selected workflow file becomes one run. - Queue. Jobs whose
needsare satisfied move frompendingtoqueuedand are published to JetStream. Postgres commits first; if the publish fails the job returns topendingand the scheduler retries. - Claim. A consumer for the job's subject claims the job atomically in Postgres (recording the owning ci process), which starts the job's clocks.
- Prepare and boot. If the job uses
vm.build, the runner's daemon builds or reuses the image. ci then creates a VM of the declared size on the chosen host. - Checkout. The guest fetches the pinned base revision from the repository URL, applies the patch, and refuses to continue unless the resulting Git tree matches the submitted tree exactly.
- Steps. Each step runs through the daemon's exec-operation API as
sh -lc. Output is masked for secrets and appended to shared log storage as it arrives; the dashboard streams it live. - Cleanup. After the last step ci captures the VM console, records the outcome, and hands the VM to durable cleanup, which stops and deletes it and only then forgets it.
- Advance. Finishing a job re-runs the scheduler for its run, queuing dependents or skipping them.
Queue subjects
Jobs are sharded so that several ci processes can consume without overlapping:
<prefix>.job.r.<runner_id> pinned to one host
<prefix>.job.n.<network_id> any eligible host in that network
<prefix>.evt.<run_id>.<job_key|run> status events (outbox publisher)
<prefix> is CI_NATS_SUBJECT_PREFIX (default ci). Two installations sharing one NATS server need different prefixes.
A running job extends its acknowledgement every 20 seconds, so ack_wait is 60 seconds regardless of job length and a dead dispatcher releases its job in about a minute. A job whose delivery fails before it is claimed is retried after 60 seconds, 5 minutes, then 15 minutes; the fourth pre-claim failure fails the job. A host's own subject is consumed one job at a time; a network subject fans out to up to four concurrent jobs across runners.
Placement
uses: on a job decides where it runs:
uses: |
Meaning |
|---|---|
| absent | The repository's assigned network, any eligible host. |
default |
The host this ci process runs on, whatever network it is in. |
<network> or <network>/* |
Any eligible online host in that network. |
<network>/<host> |
That host (by daemon ID or name). vm: builds a VM on it. |
<network>/<host>/<vm> |
An existing VM on that host. vm: is ignored, steps exec into it, and ci never stops, deletes or extends its TTL. |
For an unpinned job, ci reads each candidate daemon's GET /capabilities (to skip hosts that cannot run the job's driver) and /storage, and picks the online host with the most free disk. A host must have room for the job's data disk, twice the declared image size, and 5 GiB of headroom; a host whose free space cannot be measured is excluded. Ties go to the lower runner ID.
The network for a job without a network in uses: resolves in this order, and is stamped into the stored job plan at submit time so a redelivery runs where it was scheduled:
- the job's
uses: <network>/… - the app-lb workflow object's
network - the registered repository's assigned network (set on
/repos) - the first entry of
CI_NETWORK, or the account's default network whenCI_NETWORK=*
A submit that resolves to a network this ci instance does not serve is refused with the served list in the error.
A pinned job waits for its host. If the host is offline, the job stays queued and fails after CI_RUNNER_WAIT_SECS. Set fallback: any on the job to let it run on any host in the network instead.
VMs are disposable
Treat every job as getting a fresh VM. With vm.reuse false (or after guest filesystem corruption) ci destroys the VM as soon as the job ends. With the default vm.reuse: true it stops the VM and parks it, but the idle sweep runs with a zero-second window and destroys parked VMs on its next tick, so a later job only lands on one if it is claimed in that gap. There is no warm VM pool to plan around. What persists between jobs on a host is the image: a vm.build image is named by the hash of its Dockerfile, build context and size_mb, built once per host, and reused by every later job that names the same inputs. Unused built images are evicted later through heyvmd's ownership-checked eviction route, never while a runner has active work.
A job's data disk (disk_size_gb) is therefore new and empty in every job. Treat it as scratch space, not a cache.
Size classes
vm.size_class takes the heyvm sizes:
| Class | vCPU | Memory |
|---|---|---|
micro, mini, small |
1 | under 4 GB |
medium |
2 | 4 GB |
large |
4 | 8 GB |
xlarge |
8 | 16 GB |
Pick on memory. A microVM has no swap, so a build that needs more is OOM-killed rather than slowed down; CARGO_BUILD_JOBS in the job's env trades wall-clock time for memory headroom. ci reads back what the daemon actually allocated. A VM reported smaller than the declared class is resized once, and the job fails with both sizes named if it is still too small. heyvmd admits a VM only if the declared memory of its running VMs fits the host, so large classes are harder to place.
Native runners
A job with runs-on: instead of uses:/vm: goes to a native runner whose advertised labels include every label listed:
jobs:
mac:
runs-on: [macos, macos-intel, x86_64-apple-darwin]
steps:
- run: cargo build --release
The agent is the native-runner-agent binary built from ci/:
cargo build --locked --release --manifest-path ci/Cargo.toml --bin native-runner-agent
| Agent variable | Meaning |
|---|---|
CI_ENDPOINT |
ci base URL. HTTPS required except on loopback. |
CI_NATIVE_RUNNER_SECRET |
Must equal CI_NATIVE_RUNNER_SECRET on the ci server. |
CI_NATIVE_RUNNER_ID |
Stable runner identity. |
CI_NATIVE_PROFILE |
mac-intel or windows-x64. The agent refuses a profile that does not match the host OS and architecture. |
CI_NATIVE_RUNNER_NAME |
Display name. Default: the profile. |
CI_NATIVE_WORKDIR |
Working directory. Default: <tmp>/heyo-native. |
Profiles advertise fixed labels: mac-intel advertises namespace-profile-mac-build, macos, macos-intel, x86_64-apple-darwin; windows-x64 advertises blacksmith-2vcpu-windows-2025, windows, windows-x64, x86_64-pc-windows-msvc.
Native jobs support shell steps, conditions, env, step outputs, working directories, timeouts, continue-on-error, cancellation, ci/upload-artifact, and ci/checkout-release. Other built-in actions fail; put them in a dependent Linux job. Logs arrive with the completion report rather than streaming. Native jobs run your code directly on the host with no sandbox. The /api/native/ routes must be in app-lb public_paths; ci authenticates them itself. Protocol details are in NATIVE_RUNNERS.md.
Running it
Requirements
heyvmdon each build host, joined to a heyvm network (heyvm network add-host). The daemon must require authentication (JWT_SECRET); ci refuses a runner that answers unauthenticated.- Postgres. Migrations are compiled into the binary and applied at every startup under an advisory lock.
- NATS with JetStream enabled (
nats-server -js -sd /var/lib/ci-js). Do not use the NATS system account; JetStream cannot be enabled on it. - For image builds (
vm.build), each runner host needsdocker,mke2fs, andfakerootwhen heyvmd does not run as root.
Start
CI_HEYO_API_KEY=… \
CI_NETWORK=prod-runners \
CI_DATABASE_URL=postgres://ci:…@db/ci \
CI_WEBHOOK_SECRET=$(openssl rand -hex 32) \
CI_NATS_URL=nats://127.0.0.1:4222 \
ci
Configuration is environment-only. Every value is validated before ci binds a port or dials NATS; a bad or missing value exits non-zero with a message naming the variable (ci: refusing to start — …). The startup log prints one summary line of every resolved non-secret setting.
For a single machine with no heyvm cloud account, point ci at a local daemon:
CI_LOCAL_RUNNER=1 CI_ALLOW_UNAUTHENTICATED_RUNNERS=true CI_NETWORK=local \
CI_HEYO_API_KEY=unused CI_DATABASE_URL=postgres://…/ci CI_WEBHOOK_SECRET=0123456789abcdef ci
CI_LOCAL_RUNNER=1 uses the SDK's default local daemon URL; any other value is taken as the daemon's base URL. The runner is named local.
A supervisor program template is in ci/deploy/supervisor/ci.conf, and an app-lb static deployment for fronting the dashboard is in ci/deploy/ci.json.
Offline commands
| Command | What it does |
|---|---|
ci --check-workflows FILE… |
Parses and plans each workflow file (matrix, needs, cycles) with no configuration, database or broker. Exit 1 on any error. |
ci --inspect-executor |
Prints registered ci process boots from the database as JSON. |
ci --reconcile-service-rollout RUN_ID OPERATION_ID |
Re-checks one ci/rollout-service operation receipt using the configured database and heyosecret. |
ci --prepare-host-bootstrap …, --deliver-host-bootstrap …, --check-host-bootstrap … |
Operator tooling for the one-time native app-lb host bootstrap. See the ci README. |
Configuration
Required:
| Variable | Meaning |
|---|---|
CI_HEYO_API_KEY |
Bearer for the heyvm cloud and for each runner daemon. HEYO_API_KEY is accepted as a fallback. |
CI_NETWORK |
Networks whose hosts are runners: one name or ID, a comma-separated list (the first is the default), or * for every network on the account. |
CI_DATABASE_URL |
Postgres connection string. |
CI_WEBHOOK_SECRET |
Shared HMAC secret for submits without a repository token. At least 16 bytes. |
Service:
| Variable | Default | Meaning |
|---|---|---|
CI_NAME |
ci |
Name shown in logs and the dashboard. |
CI_LISTEN_ADDR |
127.0.0.1:9500 |
HTTP listen address. |
CI_PUBLIC_URL |
http://<listen addr> |
External base URL, used in links and submit responses. |
CI_ADMIN_EMAILS |
unset | Comma-separated emails granted admin on first sight (via app-lb forwarded identity). When unset, requests with no identity are treated as admin, which is only safe for a local, ungated install. |
CI_REQUIRE_REPO_TOKEN |
false |
Refuse the shared CI_WEBHOOK_SECRET for submits; accept only per-repository tokens. |
CI_WORKFLOW_PATH |
.ci/workflows/*.yml |
Default workflow glob. A repository registration or workflow object can override it. |
CI_MAX_SOURCE_BYTES |
67108864 (64 MiB) |
Ceiling on a submitted source descriptor after base64 decoding. |
CI_DB_STATEMENT_TIMEOUT_SECS |
30 |
Per-statement Postgres timeout. |
CI_MIGRATIONS_DIR |
unset | Extra *.sql to run after the embedded migrations. Leave unset. |
CI_LOG_DIR |
ci-logs |
Legacy local log path, imported into Postgres at startup when present. |
CI_WORKSPACE_DIR |
ci-workspaces |
Local submission staging. |
CI_LOG_RETENTION_DAYS |
2 |
Step and VM log retention. 0 keeps logs forever. Run and step rows are kept. |
CI_VM_LOG_LINES |
500 |
Lines of a VM's console captured per job. |
CI_UI_COOKIE_DOMAIN, CI_UI_COOKIE_NAME |
unset | Theme cookie scope, falling back to the fleet-wide HEYO_UI_COOKIE_* values. |
RUST_LOG |
info,ci=debug |
Log filter. |
Runners and timing:
| Variable | Default | Meaning |
|---|---|---|
CI_HEYO_BASE_URL |
SDK default | heyvm cloud base URL. |
CI_DEFAULT_NODE |
unset | Daemon ID that uses: default resolves to, overriding discovery. |
CI_DAEMON_STATE_PATH |
$MVM_DATA_DIR/daemon.json, else ~/.heyo/daemon.json |
heyvmd identity file used to resolve uses: default. |
CI_LOCAL_DAEMON_URL |
SDK default | Local daemon URL for the uses: default probe. |
CI_LOCAL_RUNNER |
unset | Drive one daemon directly instead of discovering hosts (see above). |
CI_LOCAL_RUNNER_TOKEN |
unset | Bearer for that daemon, when it requires one. |
CI_IROH_RELAY |
unset | iroh relay override for NAT traversal. |
CI_ALLOW_UNAUTHENTICATED_RUNNERS |
false |
Accept daemons that serve their API without auth (local development only). |
CI_RUNNER_REFRESH_SECS |
30 |
How often the runner list is re-read. |
CI_RUNNER_WAIT_SECS |
1800 |
How long a job may wait for a runner that is not going to take it (offline pinned host, no consumer) before it fails. Must exceed the 21-minute redelivery ladder. |
CI_QUEUE_WAIT_SECS |
unset (no cap) | Optional cap on time spent queued behind a busy runner. |
CI_MAX_JOB_SECONDS |
14400 |
Hard ceiling on any job, measured from pickup. |
CI_VM_TTL_SECONDS |
3600 |
Backstop TTL on every VM ci creates; renewed while a job runs. A job's vm.ttl_seconds overrides it. |
CI_VM_LEASE_SECS |
180 |
VM lease window; renewed every third of it. |
NATS:
| Variable | Default | Meaning |
|---|---|---|
CI_NATS_URL |
nats://127.0.0.1:4222 |
Server URL, comma-separated for a cluster, or wss://host/path through a WebSocket-capable ingress. |
CI_NATS_SUBJECT_PREFIX |
ci |
Namespace for streams, subjects and durable consumers. [A-Za-z0-9_-] only. |
CI_NATS_USER + CI_NATS_PASSWORD |
unset | User/password auth (both or neither). |
CI_NATS_TOKEN |
unset | Token auth. |
CI_NATS_CREDS |
unset | Path to a .creds file. |
CI_NATS_NKEY |
unset | nkey seed. |
Set at most one of the four NATS credential forms; naming two is a startup error. A credential embedded in the URL still works but logs a warning.
Artifacts and reports:
| Variable | Default | Meaning |
|---|---|---|
CI_ARTIFACT_SINK |
disk |
disk, s3, or artifacts (also art). |
CI_ARTIFACT_DIR |
ci-artifacts |
Directory for the disk sink. |
CI_ARTIFACT_URL |
required for artifacts |
Base URL of the artifacts store (art serve). |
CI_ARTIFACT_TOKEN |
unset | Store API key. |
CI_ARTIFACT_GUEST_URL |
CI_ARTIFACT_URL |
Store URL as reachable from inside a guest VM, when it differs. |
CI_S3_BUCKET |
required for s3 |
Bucket for the s3 sink and for debug reports. |
CI_S3_PREFIX |
ci |
Key prefix. |
CI_S3_REGION, CI_S3_ENDPOINT |
unset | S3 region and endpoint. Credentials come from the standard AWS chain. |
When CI_S3_BUCKET is set, ci also uploads a per-job debug report (job identity, outcome, step logs, VM console) to <prefix>/<run>/<job>/debug-<attempt>-<sandbox>.json before deleting the VM, regardless of the artifact sink. Use a private bucket.
Integrations:
| Variable | Default | Meaning |
|---|---|---|
CI_HEYOSECRET_URL |
HEYOSECRET_URL |
heyosecret base URL. |
CI_HEYOSECRET_TOKEN |
HEYOSECRET_INTERNAL_API_KEY, then PLATFORM_INTERNAL_API_KEY |
heyosecret bearer. Never reaches a build. |
CI_APP_LB_URL, CI_APP_LB_TOKEN |
unset | app-lb admin API for workflow objects. Without it, ci uses repository registrations and CI_WORKFLOW_PATH only. |
CI_NATIVE_RUNNER_SECRET |
unset | Bearer for /api/native/*. Native runners are disabled without it. |
Release, deployment and host-maintenance variables (CI_RELEASE_POLICIES, CI_HOST_APP_LB_TARGETS, CI_HOST_MAINTENANCE_TARGETS, CI_HOST_HEYVM_BOOTSTRAP_TARGETS, CI_CONTROLLER_*, CI_APPLICATION_*, CI_EXPECTED_SHA) are operator configuration for the built-in release actions. They are documented with those actions in the ci README.
Fronting it with app-lb
app-lb's sign-in gate admits browsers only, so the machine routes must be listed in the deployment's public_paths; each authenticates its own requests:
"public_paths": [
{ "path": "/healthz", "scope": "public" },
{ "path": "/api/submit", "scope": "public" },
{ "path": "/api/native/", "scope": "public" },
{ "path": "/api/stream/", "scope": "public" },
{ "path": "/api/runs/", "scope": "public" },
{ "path": "/__ui/", "scope": "public" }
]
Keep /repos, /maintenance, /networks/*/join and the /vms actions behind the gate. ci reads the forwarded identity headers (x-auth-request-user, -email, -name), which are trustworthy only on a gated deployment, and keeps its own admin list seeded from CI_ADMIN_EMAILS. See app-lb auth.
Submitting with git submit
Install and configure
ci/install-git-submit.sh # installs git-submit into ~/.local/bin
Register the repository on the dashboard's /repos page and mint a token. The page shows the token once, with the two lines to run in your clone:
git config ci.endpoint https://ci.us2.heyo.work
git config ci.token cis_…
| Setting | Environment override | Meaning |
|---|---|---|
ci.endpoint |
CI_ENDPOINT |
ci base URL (or the full /api/submit URL). |
ci.token |
CI_TOKEN |
Repository token, sent as a bearer. |
ci.secret |
CI_WEBHOOK_SECRET |
Shared secret; the payload is HMAC-signed (x-heyo-signature-256) when no token is set. |
ci.workflowPath |
Extra workflow globs (repeatable) when the repository uses a custom path. |
CI_SUBMIT_TIMEOUT (default 300 seconds) bounds the upload. The client needs git, curl, base64, python3, and openssl for the shared-secret path.
Usage
git submit # submit HEAD
git submit --dry-run # show what would be sent
git submit --dirty # include uncommitted tracked changes
git submit --ref <rev> # submit another commit
git submit pr59 # submit the exact head of GitHub PR #59
git submit --only app-lb # run only this workflow (repeatable)
git submit --workflow <id> # attribute the run to a named workflow object
git submit --submit-empty # submit even when the tree equals the default branch
git submit --version
git submit --upgrade # reinstall from the release channel
The client prints one run URL per started workflow, plus a submission URL when a release workflow was selected, and any warnings (for example, a pinned host that is offline).
What gets sent:
- A pinned base and a patch. The client asks
originwhich heads it advertises, picks the nearest published ancestor of the revision you are submitting, and sends a binary full-index patch from that base to your target tree. A pushed commit is its own base with an empty patch. Local commits and--dirtywork as long as some ancestor is published; an unpublished root-only repository cannot be submitted.--archivehas been removed. - Workflow YAML from the target tree (not your worktree):
.ci/workflows/*.ymland*.yamlby default, at most 256 KiB per file and 1 MiB total. - The changed-path list for the full diff from the default branch, which drives
paths:filters andchanged(). If the diff cannot be computed (root commit, base unavailable, more than 5000 paths), the change set is unknown and every path filter matches.
The runner, not ci, does the checkout: it fetches the base from the repository URL, applies the patch, and verifies the tree. Private repositories therefore need a checkout credential in heyosecret (CI_GIT_AUTH_TOKEN or GITHUB_TOKEN in the workflow's secret scope). The submit token is not repository read access.
pr<number> fetches refs/pull/<number>/head into a temporary object store without touching your branch, index or worktree. It cannot be combined with --ref or --dirty.
--only <selector> runs just the named workflow files. A selector is the file path, its basename with or without extension, or the workflow's name:, matched case-insensitively. A named workflow runs even if its branches/paths filters would have declined, and its run is created with an unknown change set so job-level changed() conditions also pass. A selector that matches nothing fails the submit.
Repository tokens
A repository token is minted per registration, stored as a SHA-256 digest, shown once, and revocable on its own. A submit whose payload names a different repository than its token is refused. Pausing a registration refuses its tokens without deleting them. CI_WEBHOOK_SECRET works for any repository and cannot be revoked per repository; set CI_REQUIRE_REPO_TOKEN=true once every repository has a token.
A registration can also set a workflow glob override and an assigned network.
Workflow objects
With CI_APP_LB_URL set, ci polls app-lb for workflow objects, each pointing at a repository, a ref, a path glob and a network:
heyctl create workflow build \
--repo git@github.com:me/app.git \
--network prod-runners \
--path '.ci/workflows/*.yml'
heyctl get workflows
Objects match a submit by repository URL (git@github.com:me/app.git and https://github.com/me/app are the same). Several objects may name one repository; each produces its own runs. There is no in-place edit; delete and re-create with the same ID. See heyctl.
Workflow reference
name: build
on:
submit:
paths: ['api/**', 'Cargo.lock']
jobs:
build:
uses: prod-runners # any eligible host in this network
vm:
driver: firecracker
build:
dockerfile: .ci/image/Dockerfile
size_class: large
disk_size_gb: 20
strategy:
matrix:
target: [x86_64, aarch64]
env:
CI_ENVIRONMENT: prod # selects the secret scope
timeout-minutes: 75
outputs:
version: ${{ steps.stamp.outputs.version }}
steps:
- id: stamp
run: echo "version=$(git describe --always)" >> "$CI_OUTPUT"
- name: Build
working-directory: /workspace/api
run: cargo build --release --locked --target ${{ matrix.target }}
timeout-minutes: 60
env:
DATABASE_URL: ${{ secrets.DATABASE_URL }}
- uses: ci/upload-artifact
with:
name: api-${{ matrix.target }}
path: dist
deploy:
needs: [build]
if: ${{ needs.build.result == 'success' && ci.branch == 'main' }}
steps:
- run: echo "deploying ${{ needs.build.outputs.version }}"
Every level rejects unknown keys (stpes:, timeout_minutes:) with a parse error naming the job, except inside vm:, where an unknown key is currently ignored. Run ci --check-workflows before you submit.
Top level
| Key | Meaning |
|---|---|
name |
Display name, also usable as an --only selector. |
on |
submit (default when absent), a list, or a mapping with a submit: filter block. release marks the repository's single release coordinator (see Releases). Other triggers parse but never start a run. |
env |
Accepted, but not currently applied to steps. Put variables on the job or step. |
jobs |
Ordered mapping of job ID to job. IDs become queue subject tokens, so keep them to letters, digits, - and _. |
on: submit filters
on:
submit:
branches: [main, 'release/*'] # or branches-ignore
paths: ['packages/api/**'] # or paths-ignore
| Pattern | Matches |
|---|---|
* |
Any characters within one path segment. |
** |
Any number of whole segments, including none. |
? |
One character within a segment. |
paths builds when any changed path matches; paths-ignore skips only when every changed path matches. branches and branches-ignore cannot be combined, nor can paths and paths-ignore, and a leading ! is refused. An unknown change set matches every filter. A submit where every workflow declines is a success with no runs, and git submit prints each reason.
Job keys
| Key | Meaning |
|---|---|
name |
Display name. |
uses |
Placement (see Placement). |
runs-on |
Native runner labels. Mutually exclusive with uses and vm. |
fallback |
none (default) or any: whether a pinned job may run on another host in the network. |
vm |
Machine to boot (below). |
needs |
Job IDs this job waits for. Cycles and unknown IDs are parse errors. |
if |
Condition. Without one, the job runs only if every dependency succeeded; failure, cancellation and skips propagate down the chain. |
strategy.matrix |
Axes of values, plus optional include and exclude lists. |
strategy.max-parallel |
Cells that may run at once. |
strategy.fail-fast |
Default true. |
env |
Step environment for every step in the job. Matrix expressions are substituted at plan time. CI_ENVIRONMENT here selects the secret scope. |
outputs |
Job outputs, readable as needs.<job>.outputs.<name>. |
continue-on-error |
Job failure does not fail the run. |
timeout-minutes |
Default 60. |
steps |
Required, non-empty. |
vm:
| Key | Meaning |
|---|---|
driver |
Required when vm: is present. firecracker (the default when vm: is absent) or kvm. libvirt is refused. |
image |
A base image name the host already has, for example ubuntu:24.04. |
build.dockerfile |
Dockerfile path in the submitted tree. Mutually exclusive with image. |
build.context |
Build context. Default: the Dockerfile's directory. |
build.size_mb |
Rootfs size. Default: auto-sized from the exported tree (about 1.2x plus 64 MB). |
size_class |
micro, mini, small, medium, large, xlarge. |
disk_size_gb |
Extra data disk for the job. |
working_directory |
Default working directory for steps. Default /workspace. |
env_vars |
Environment passed to the VM at creation. |
setup_hooks |
Commands run when the VM is created. |
ttl_seconds |
Create-time VM TTL. Set it above the job timeout for jobs longer than CI_VM_TTL_SECONDS. |
cache_key_files |
Feeds the VM fingerprint and triggers source preparation. Parked VMs are swept almost immediately, so it rarely leads to reuse. |
reuse |
Default true: park the stopped VM instead of destroying it. The idle sweep retires parked VMs on its next tick, so reuse is incidental. |
vm.build runs docker build, docker export, then mke2fs on the runner, so Dockerfile features all work but OCI metadata does not survive: ENV, CMD and ENTRYPOINT are dropped. Put environment a step needs in /etc/profile.d with a RUN (steps run under sh -lc, a login shell). The image must have an /init.sh that prints HEYVM_READY; .ci/image/ci/init.sh is a working example. The first job on a host that needs an image builds it and later jobs wait for that one build.
Steps
| Key | Meaning |
|---|---|
name, id |
Label, and the ID used in steps.<id>.outputs. |
run |
Shell script. Exactly one of run or uses. |
uses |
A built-in action (below). Repository composite actions are not supported. |
with |
Action inputs. Scalars only; true and 3 are read as strings. |
if |
Step condition. |
env |
Step environment, overriding the job's. |
working-directory |
Used verbatim as cd <value>, so give an absolute path such as /workspace/api. Default: the job's working directory (/workspace). |
timeout-minutes |
Default 30. |
continue-on-error |
Step failure does not fail the job. |
Every step gets CI=true, CI_JOB (job ID) and CI_JOB_KEY (matrix-expanded key). Write name=value lines to $CI_OUTPUT to set step outputs. The checkout lands in /workspace and is wiped on every checkout; paths in ci/upload-artifact and cache_key_files are relative to the repository root.
Expressions
${{ }} is substituted in run, env, with, if and job outputs. Contexts:
| Context | Contents |
|---|---|
ci |
sha, before, ref, branch, repository, run_id, workflow, release_base_sha, changed_files (array), changes_known (bool), changes_reason (empty when known). |
matrix |
The current cell's values. |
env |
The job's env. |
job |
id, key. |
needs |
needs.<job>.result (success, failure, cancelled, skipped; a matrix collapses to its worst cell) and needs.<job>.outputs.*. |
steps |
steps.<id>.outputs.*. |
secrets, vars |
From heyosecret (below). |
Functions: contains, startsWith, endsWith, format, join, toJSON, fromJSON, success, failure, cancelled, always, and changed(pattern, …).
changed() uses the same glob matcher as paths: and is true when any changed file matches any argument, or when the change set is unknown. Use it to gate jobs in a monorepo:
jobs:
api:
if: ${{ changed('packages/api/**', 'Cargo.lock') }}
An explicit if: replaces the default "all dependencies succeeded" rule, which is how if: ${{ always() }} runs a cleanup job after a failure.
Secrets and variables
ci resolves secrets once per job from heyosecret under the prefix:
ci/<workflow>/<environment>/<NAME>
<workflow>is the app-lb workflow object's ID when one matched; otherwise the registered repository's name (by default the last segment of its URL); otherwise the submitted repository's directory name. It is exposed as${{ ci.workflow }}.<environment>is the job'senv.CI_ENVIRONMENT, defaulting todefault.- Characters outside
[A-Za-z0-9-_.:@]are replaced with-.
An entry tagged public in heyosecret becomes vars.<NAME> and is shown in plain text. Everything else becomes secrets.<NAME> and is masked (***) in logs before they are stored or streamed; values shorter than four characters are not masked. The heyosecret token itself never reaches a build. The --secrets-prefix field on a workflow object is stored but not currently applied.
Store the checkout credential for a private repository as CI_GIT_AUTH_TOKEN (or GITHUB_TOKEN) in the same scope.
Artifacts
Uploading
- uses: ci/upload-artifact
with:
name: app # required
path: dist # required, relative to the repository root
description: App binary for x86_64-linux # artifacts sink only
public: true # artifacts sink only
alias: app-live # artifacts sink only
The step packs path into a tar.gz in the guest. With the artifacts sink, the guest pushes the tarball straight to the store with curl -T (using CI_ARTIFACT_GUEST_URL or CI_ARTIFACT_URL); if that fails, ci reads it out through the exec channel in chunks, which is much slower on Firecracker. Both directions are checked by SHA-256.
| Sink | Where it goes | Notes |
|---|---|---|
disk |
CI_ARTIFACT_DIR on the ci host |
Not shared between ci instances. |
s3 |
CI_S3_BUCKET / CI_S3_PREFIX |
No public links. |
artifacts |
The artifacts store at CI_ARTIFACT_URL |
Supports description, public and alias. |
In the artifacts store each upload is tagged:
ci-<workflow>-<run>-<job_key>-<name>
truncated to 64 characters. Run IDs are fixed-width hex (%012x-%08x, time then sequence), so sorting these tags lexicographically sorts them by time.
descriptionis stored as a label beside the digest, shown in the store and by installers'--list.public: truemarks the blob public (PUT /public/{digest}), so{store}/blobs/{digest}downloads with no credential. The link is printed in the step log as[ci] public link: …and shown on the run page. Tags and listings still need the key.aliasmoves a second, stable tag onto this upload, so a deployment can follow the newest green build by name. Aliases cannot start withci-, and a failure to set one fails the step.
Each upload writes a ci.artifact.published.v1 event. A published artifact is not a release approval: a later step can still fail.
Downloading in a later job
jobs:
package:
needs: [build]
steps:
- uses: ci/download-artifact
with:
name: app
path: app.tar.gz # the uploaded tar.gz, not unpacked
job: build # only when several producers used this name
- run: tar -xzf app.tar.gz
The producer must be in needs and must have succeeded. Size and SHA-256 are verified. Cross-job files must go through artifacts; each job's VM is deleted when it finishes. In a coordinated release, with.workflow names a validation workflow path from the same submission.
Dashboard and HTTP API
Pages
| Path | Shows |
|---|---|
/ |
Recent runs (latest 50). |
/runs/{run_id} |
Jobs, queued time and duration, artifacts and public links, VM logs, Release, Deployments, and an event timeline. Admin buttons: Cancel, Run again, Re-run failed jobs. |
/runs/{run_id}/jobs/{job_key} |
One job's steps and live log. |
/networks (also /runners) |
Every heyvm network on the account with its hosts, which ones this instance serves, each host's and network's queue depth read from JetStream, and daemons that joined no network. Admins can join this host to a network. |
/vms |
ci-owned VMs by host, including ones still being created (building), with declared vs. reported size. Admin actions: destroy, resize, clean up failed. |
/workflows |
app-lb workflow objects. |
/repos |
Admin: register repositories, mint and revoke tokens, pause, assign a network. |
A no consumer flag on an online host in the Queue column means jobs are routed to a subject nothing reads: the host is not in a network this instance serves, or its consumer is not bound.
Cancelling marks the run and all unfinished jobs cancelled. A running step cannot be aborted in the guest; ci stops waiting within seconds, records it cancelled, and cleans up the VM while the command finishes or times out. A re-run is always a new run (rerun_of links them). Re-run failed jobs carries successful jobs and their outputs over and schedules the rest.
Machine API
These routes authenticate with the repository token (Authorization: Bearer cis_…) scoped to the run's repository, or with an HMAC of the request path signed by CI_WEBHOOK_SECRET. A run from another repository answers 404.
| Method and path | Purpose |
|---|---|
POST /api/submit |
Submit (used by git submit). Accepts the bearer token, or an x-heyo-signature-256: sha256=<hex> HMAC of the body. Returns 202 with runs, url, warnings, and submission when a release coordinator was created. |
GET /api/runs/{run_id} |
Run status, jobs, steps, reruns, debug_reports. |
GET /api/runs/{run_id}/logs?job=&tail=&failed_only= |
Step logs, tail-first. tail defaults to 16384 bytes per step, capped at 1 MiB. Includes the checkout (-1) and VM console (-2) pseudo-steps. |
GET /api/runs/{run_id}/events?limit=&before= |
Durable status events, newest first. limit defaults to 50, max 100. |
GET /api/runs/{run_id}/deployments |
Deployment operations started by the run. |
GET /api/runs/{run_id}/release |
Release publication state (prepared, uncertain, confirmed) and SHAs. |
POST /api/runs/{run_id}/rerun-failed |
Re-run failed jobs. Bearer token only; 409 while the run is active or a deployment is unresolved. Each accepted call creates a new run. |
POST /api/runs/{run_id}/cache/{sandbox_id}/destroy |
Destroy an idle VM last used by this run. Bearer token only. |
GET /api/stream/{run_id}/{job_key} |
Server-sent log stream, authorized by a short-lived token the run page mints. |
GET /healthz |
ok, with x-ci-revision and x-ci-binary-sha256 headers when known. Unauthenticated. |
The /api/native/* routes serve native runners (bearer CI_NATIVE_RUNNER_SECRET), and /api/lifecycle/* serves managed-platform lifecycle calls (bearer CI_APPLICATION_LIFECYCLE_TOKEN).
curl -fsS -H "Authorization: Bearer $(git config ci.token)" \
"https://ci.us2.heyo.work/api/runs/$RUN/logs?failed_only=true&tail=4096"
Multiple instances, drain and maintenance
Several ci processes can serve the same installation, including one per region, as long as they share the same Postgres database and the same NATS server and subject prefix. Every registered process may execute work; there is no leader. A job is claimed atomically by one process, which owns its VM and cleanup until the cleanup is confirmed. If a process dies, its unclaimed queue messages are redelivered to another; work it had already claimed is not taken over automatically and needs reconciliation.
Maintenance routes are admin-only (behind the app-lb gate, never in public_paths), and each request applies only to the process that receives it. Send x-ci-target-boot: <boot id> to have any other process reject the request with 409. ci --inspect-executor lists boot IDs.
Draining one ci process
POST /maintenance/{operation-uuid}/pause stop admitting new jobs on this process
GET /maintenance this process's outstanding work
POST /maintenance/{operation-uuid}/quiesce succeeds once its jobs and cleanups are done
POST /maintenance/{operation-uuid}/resume reopen admission for the same operation
Unclaimed deliveries go back to NATS for the other process to take. Running jobs finish and clean up normally.
Draining a runner host
Pausing a ci process does not stop the other process from placing jobs on a host. To empty a host, use the runner drain on either process:
POST /maintenance/runners/{runner-id}/{operation-uuid}/pause
GET /maintenance/runners/{runner-id} running jobs, host work, cleanup, drained
POST /maintenance/runners/{runner-id}/{operation-uuid}/resume
Use the canonical runner ID, not its display name. The drain is stored in Postgres and survives restarts. Jobs already running finish on that host. New and queued unpinned jobs go to another healthy host in the same network; jobs pinned with uses: <network>/<host> (without fallback: any) and jobs targeting an existing VM stay pinned and wait. The host is drained when admission is closed and running jobs, outstanding host work and cleanups are all zero. Every ci process sharing the host must run a version that implements runner drain.
Releases and deployments
A repository may define one workflow with on: release. On a full git submit, ci creates the selected validation runs and a release coordinator run together; the coordinator waits until every selected validation succeeds and then runs its jobs. --only, reruns and partial selections never run the coordinator. The coordinator must start with a merge job containing only ci/merge-release, and every deployment job must depend on it.
Built-in release and deployment actions:
| Action | Purpose |
|---|---|
ci/merge-release |
Fast-forward the default branch to the validated commit (optionally with version bumps and tags). |
ci/checkout-release |
Check out the confirmed release commit. |
ci/download-artifact |
Fetch an artifact from this run or, with workflow, from this submission's validation runs. |
ci/publish-service-archive, ci/promote-service-archive |
Upload a packaged tarball as an Orchestrator service archive. |
ci/deploy-service |
Start an Orchestrator service deployment. |
ci/rollout-service |
Candidate-first rollout of an app-lb deployment from a validated artifact. |
ci/publish-rootfs, ci/deploy-app-lb |
Publish a rootfs image and roll it out to an app-lb deployment. |
ci/rollout-host-app-lb |
Replace the app-lb executable on a host. |
ci/rollout-host-heyvmd, ci/bootstrap-host-heyvm, ci/host-heyvm-maintenance |
Host daemon maintenance. |
ci/deploy-controller |
Replace the ci controller itself; must be last. |
Sequence regional stages with needs so a failure in one region stops the next. Deployment actions record an operation ID before calling the remote API and reconcile that same operation after a restart or lost response rather than submitting another. Their inputs, the operator-owned target mappings they require, and recovery endpoints are specified in the ci README. For the platform-level view of regional rollouts, see multi-region and orchestrator.
This repository's release pipeline
Each installable component has one workflow file in .ci/workflows/, with one job named release that builds, tests, and uploads a tarball with ci/upload-artifact:
| Workflow | Artifact | Contents |
|---|---|---|
app-lb.yml |
app-lb |
app-lb, heyctl, app-lb.conf |
app-obs.yml |
app-obs |
app-obs, app-obs-dump, app-obs.conf |
ci.yml |
ci |
ci, ci.conf, migrations/ |
art.yml |
art |
art |
queue.yml |
queue |
queue, queue.conf, a deployment template |
heyosecret.yml |
heyosecret |
heyosecret, heyosecret.conf, migrations/ |
pg-fc.yml |
pg-fc |
pg-vm-pool |
codegraph.yml |
codegraph |
codegraph |
orchestrator.yml |
orchestrator-linux |
Orchestrator release bundle |
Every artifact also carries REVISION or BUILD-INFO and a SHA256SUMS file. Each workflow's paths: filter lists its crate, its path dependencies, its own workflow file and its image directory under .ci/image/, so a commit only rebuilds what it touches. regional-release.yml is the on: release coordinator that merges and then rolls validated artifacts out to us3, then eu1, then replaces the ci controller.
Publishing for installation is a separate step run by someone with the store's API key:
- The workflows upload to the artifacts store under
ci-<workflow>-<run>-release-<name>tags. .ci/publish-releases.shresolves each app's newest tag to a blob digest, marks that blob public, and writes one<app>.jsonmanifest per app plusindex.json,install-apps.shandbootstrap-host.sh.--push-tag releases-liveuploads that directory as one public bundle and moves the tag; thereleasessite deployment serves it atget.us2.heyo.work.
ART_API_KEY=… sh .ci/publish-releases.sh --all \
--from-url https://get.us2.heyo.work --push-tag releases-live
heyctl pull releases
The installer then downloads public blobs with no credential and verifies the digest and SHA256SUMS. To install the binaries, see installation.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
ci: refusing to start — CI_… … |
A required variable is missing or invalid. The message names it. |
| Submit returns 401 | Missing or wrong ci.token/ci.secret, or /api/submit is not in app-lb public_paths (the gate answers 401 to non-browser clients). |
| Submit refused: network not served | The job's resolved network is not in CI_NETWORK. Add it, use *, or reassign the repository's network. |
| Submit says a selector matched no workflow | --only value does not match a file path, basename or name:; or the file lacks submit in on:. |
| Submit started no runs | Every workflow's filters declined. git submit prints each reason; --only bypasses filters. |
| Every path filter matched | The change set was unknown (no published ancestor diff, root commit, more than 5000 paths). Check ${{ ci.changes_reason }}. |
Job stuck queued, host shows no consumer on /networks |
The job is pinned to a host that is offline or not in a served network. It fails after CI_RUNNER_WAIT_SECS; set fallback: any or fix the host. |
| Job fails with "no runner took this job within …s", "no host in network … is online", or "… free disk bytes … this job requires …" | No eligible host: none online, none supports the job's driver, or none has enough free disk (data disk + 2x image size + 5 GiB). The error names the cause; check /networks and host disk. |
| Runner refused as unauthenticated | The daemon serves its API without auth. Configure JWT_SECRET on heyvmd, or set CI_ALLOW_UNAUTHENTICATED_RUNNERS=true for local use only. |
| Checkout fails on a private repository | No CI_GIT_AUTH_TOKEN/GITHUB_TOKEN in ci/<workflow>/<environment>. |
| Checkout fails with a tree mismatch | The base commit or patch does not reproduce the submitted tree. Resubmit after git fetch. |
| Image builds but every boot fails | The image lacks /init.sh printing HEYVM_READY. |
A variable set with ENV in the Dockerfile is missing |
OCI metadata is dropped by docker export. Write it to /etc/profile.d in a RUN. |
cd: no such file at the start of a step |
working-directory is used verbatim. Use an absolute path under /workspace. |
| Step killed at 30 minutes | Default step timeout. Set timeout-minutes on the step; the job default is 60. |
| Long job's VM disappears mid-build | VM TTL shorter than the job. Set vm.ttl_seconds above the job's timeout. |
| Build OOM-killed, or job fails "VM too small" | Raise size_class or lower CARGO_BUILD_JOBS. The daemon may also have allocated less than declared; see the size on /vms. |
| "host memory capacity unavailable" | heyvmd admission counts declared memory of running VMs. Free capacity or add a runner host. |
uses: ci/… is not a built-in action |
Only the actions listed above exist; repository composite actions are not supported. |
vm: setting has no effect |
Unknown keys inside vm: are ignored; check spelling with the table above. |
| Artifact upload is very slow | The guest could not reach the store and fell back to the serial exec channel. Set CI_ARTIFACT_GUEST_URL to a URL the guest can reach; the step log explains why the push failed. |
| Public link missing | public works only with CI_ARTIFACT_SINK=artifacts. |
| Logs of an old run are empty | Logs are deleted after CI_LOG_RETENTION_DAYS (default 2); step rows remain. |
| Log stream stops after a ci restart | Stream tokens are per process. Reload the run page. |
| NATS login succeeds, then the first stream fails | You are using the NATS system account. Use a regular account with JetStream enabled. |
| Two installations steal each other's jobs | They share a NATS server and prefix. Give each a distinct CI_NATS_SUBJECT_PREFIX. |
For offline validation of every workflow file:
ci --check-workflows .ci/workflows/*.yml
This checks parsing and planning only, not that the build commands work.