skip to content
terminalhire documentation

Verifying your work — th run

th run takes the diff sitting in your working tree, applies it to a clean clone of the target repository, and runs the project's own install and test commands inside a container. You get a verdict, and a preview URL you can open to read the whole run.

It is for the gap between "the tests pass on my machine" and "I submitted it and waited." You find out before the poster does.

What you need

Docker running, and network access for the first clone of the target. Nothing else — the engine ships inside the CLI.

The short version

From inside the workspace claim start (or claim slice) delivered for a poster's task:

th run

The claim ledger knows which claim the workspace belongs to and which commit it was delivered from, so it fills in --claim and --sha. The diff runs from that commit to your working tree, so work you have already committed is tested too: the same change claim submit sends.

  • A whole-repository delivery is its own target. th run clones the delivered commit out of the workspace itself, so a private repository needs no credential.
  • A delivery of only the files you were granted needs the clone URL of the full repository: th run --target <url>. If you don't have access to it, skip the local run. The poster-side run after claim submit still happens, unless the poster reviews the work by hand; see Post work.

Anywhere else, th run tests only your uncommitted changes, and you pass all three yourself, as flags or once in .th-run.json:

th run --claim <id> --target https://github.com/koajs/koa.git --sha <40-hex>

Run it again and again while you work; add --watch to re-run on every save.

Defaults live in .th-run.json in your checkout, so a claim you run repeatedly needs no flags at all. Every flag overrides the file.

Inside a claim workspace, a run tests your change against the commit you were delivered. When you submit, the poster-side run puts your change on top of the repository's default branch as that branch stands then, on a branch of its own for the poster to review. If the default branch has moved since delivery, the two results can differ. A run inside a claim workspace prints a line reminding you of this, unless you pass --json.

Options

Flag What it does
--claim <id> The claim this work belongs to. Inside a claim workspace, read from the claim ledger
--target <git-url> Repository the work is verified against. A whole-repository delivery supplies its own
--sha <40-hex> The commit your diff applies on top of. Inside a claim workspace, the commit it was delivered from
--slice <a,b,c> Optional local guard rail: files you mean to touch. A diff touching anything else is refused before any container starts, and nothing re-derives that list at submission
--local <dir> Checkout to read the working diff from (default: the current directory)
--watch Re-run whenever a file in the checkout changes
--keep <seconds> Hold the preview open this long (default: until you press Ctrl-C)
--no-preview Skip the preview URL
--no-screenshots Skip the before/after screenshots. See Screenshots
--preview-route <r> Routes to screenshot, comma-separated (default: /; at most 5)
--json Print the result as JSON instead of a report
--test-command <cmd> Override the derived test command. The override is recorded in the result
--placement <kind> Where the run happens: local-docker (default) or hosted. See Where it runs

How a run is fenced

A run is two steps, and they are fenced differently on purpose.

Install gets a network, through one door. Dependencies have to come from somewhere, so the install step reaches the outside only through an allowlisting proxy on its own internal network. There is no route around it.

A lockfile that falls behind is installed around, and named. If a claim adds or changes a dependency without including the lockfile, package.json asks for entries the lockfile does not list. A frozen install (npm ci, yarn install --frozen-lockfile, pnpm install --frozen-lockfile) would refuse that outright. When the run finds that gap, it installs without the frozen check and prints which entries are missing. The run can still pass; the line says to include the updated lockfile in the patch, or regenerate it before merging. A lockfile that lists everything stays frozen.

Tests get no network at all. The test step runs with networking switched off entirely — no interface, no proxy, nothing to reach. A test suite that quietly depends on a live service fails here, which is the point: so does the poster's CI.

A posting can set its own test time limit. When the posting you claimed has one, measured by running its tests on Terminal Hire's machines, th run stops the test step at that limit (5 to 120 minutes) instead of the default. The report names it, as in time limit 45 min, with the install and test steps' own times.

The container drops all capabilities, forbids privilege escalation, and caps process count. The repository's own code never runs anywhere else.

The preview

Unless you pass --no-preview, a local run ends with a URL. Opening it shows the same result the terminal printed — the verdict, the test output tail, and what was actually run.

The preview binds to loopback only. It is reachable from your machine and nowhere else, and it stops when you press Ctrl-C or when --keep runs out. Expiry and revocation beyond that are not built yet.

The preview never changes the verdict. If it cannot start, the run says why on a preview not started: line and returns its result as if you had passed --no-preview. A hosted run always takes that path for now: serving a preview from the venue would need ingress to it, which is not built.

Screenshots

When the repository can be built or served, a local run ends with before and after screenshots. Once the tests have a verdict, the run builds your version of the tree with no network, takes one screenshot of each route at desktop size in light mode, and then does the same for the tree without your change. The PNGs are saved under ~/.terminalhire/screenshots/, and the report lists where:

screenshots  2 in ~/.terminalhire/screenshots/2026-09-24T18-08-41-314Z-1w4zbq (routes /; light; desktop)
             after: 1 shot
             before: 1 shot
             rendered with no network: no backend, no signed-in state

Which repositories qualify is read from what the project declares, not from a list of frameworks: a build script, a preview, start, serve or dev script, or an index.html. If the build writes an index.html, that output is served as static files. Otherwise the run starts the serve script, again with no network. A repository that qualifies on neither count is skipped, and the report says why in one line, for example screenshots skipped — nothing here can be served. A build that needs the network, such as one that downloads fonts, is skipped the same way. So is a workspace root.

The pictures show what the app renders on its own: nothing it would fetch from a backend, and no signed-in state.

Screenshots never change the test result, and they stay on your machine. The first run builds the screenshot image, about 900 MB. --watch never takes screenshots. --json output carries them as a screenshots object with each file's path and SHA-256.

Your own screenshots

On a first-party posting you can add pictures you took yourself to the submission. Take them with Playwright, or reuse the ones the run saved, list them in a small JSON file, and pass it to terminalhire claim submit <id> --screenshots <file>:

{
  "v": 1,
  "shots": [
    { "file": "./shots/home-dark.png", "caption": "Home in dark mode" },
    { "file": "./shots/settings.png", "caption": "Settings, saved state" }
  ]
}

Paths resolve against the JSON file's directory. Up to six pictures, each a PNG of at most 4 MiB (4,194,304 bytes) and at most 8192 pixels on a side, with a caption of up to 140 characters. The CLI checks all of that before it asks you to confirm, and a bad file is refused with nothing sent. The pictures upload after the patch is applied. One that does not arrive never fails the submission; the claim page lists it as not arrived.

The poster sees them on the claim page beside the verification run's own pictures, under a heading that says they came from you. Send pictures of this patch as it runs, never edited ones.

Where it runs

By default a run happens on your machine, in your Docker daemon. That is what --placement local-docker means.

--placement hosted asks for a hosted pool: a fresh cloud VM booted for your run and torn down after it — teardown is attempted per run, and the platform deletes the instance at its one-hour lifetime either way — so the verification happens on a machine neither you nor the poster operates. It needs cloud credentials on the machine that launches it, so it is meant for our own dispatch worker rather than a developer's terminal — a bare th run --placement hosted without those credentials refuses before any machine is booted, and nothing about your work leaves your machine when it does. It refuses the same way on Windows, whatever credentials are present: the hosted machine is reached through a forwarded Unix socket, which OpenSSH on Windows cannot forward. A run's record says what it can prove: a hosted record our intake accepted carries a Google-signed machine identity the intake checked, a run on the worker's own machine carries only what the venue reported about itself, and an absent identity proves nothing about where a run happened — a hosted run that failed partway, or whose result never reached the intake, leaves none.

A refusal exits 2, not 1. In this command 1 means your tests ran and failed; 2 means we declined to do something, such as run them at all. A preview that cannot start is never a refusal: the run still returns its verdict. A script can tell the two apart, and nothing we refuse to do will ever be reported as a fault in your work.

What a result is, and is not

A run tells you what the project's own tests did, on a clean tree, with your diff applied. That is useful and it is checkable.

It is not an attestation that anything ran on hardware anyone can vouch for. The result records this about itself in plain terms rather than implying otherwise. A hosted run dispatched by our worker can go one step further — when its record lands, it carries the venue VM's Google-signed instance identity, checked at intake — but that names the machine, not the code it ran, and the result keeps saying which claim it is making.

Where this fits

th run is the step before claim submit. Run it until it is green, then submit. On a first-party posting you can read the poster's own verdict afterwards with claim runs <id>.

If the run on your submission comes back red

Send another one. A fix is a new commit, and a new commit gets its own run — you are not limited to one submission, and you never were.

Every run on a claim is kept, and the page lists them newest first — up to the twenty most recent — each against the commit it tested. A run is numbered from your first, so a capped list never tells you the eleventh try was the first. The poster's own page lists the same runs. Three reds above a green is not a mark against the work; it is how the work got there.

A run that was refused is different from one that failed. Refused means we never got as far as testing your code, so it says nothing about it — you can send a newer commit, but there may be nothing to fix.