Verifying your work — th run
th run takes the diff sitting in your working tree, applies it to a clean
clone of the target repository, and runs the project's own install and test
commands inside a container. You get a verdict, and a preview URL you can open
to read the whole run.
It is for the gap between "the tests pass on my machine" and "I submitted it and waited." You find out before the poster does.
What you need
Docker running, and network access for the first clone of the target. Nothing else — the engine ships inside the CLI.
The short version
From inside the workspace claim start (or claim slice) delivered for a
poster's task:
th run
The claim ledger knows which claim the workspace belongs to and which commit it
was delivered from, so it fills in --claim and --sha. The diff runs from that
commit to your working tree, so work you have already committed is tested too:
the same change claim submit sends.
- A whole-repository delivery is its own target.
th runclones the delivered commit out of the workspace itself, so a private repository needs no credential. - A delivery of only the files you were granted needs the clone URL of the
full repository:
th run --target <url>. If you don't have access to it, skip the local run. The poster-side run afterclaim submitstill happens, unless the poster reviews the work by hand; see Post work.
Anywhere else, th run tests only your uncommitted changes, and you pass all
three yourself, as flags or once in .th-run.json:
th run --claim <id> --target https://github.com/koajs/koa.git --sha <40-hex>
Run it again and again while you work; add --watch to re-run on every save.
Defaults live in .th-run.json in your checkout, so a claim you run repeatedly
needs no flags at all. Every flag overrides the file.
Inside a claim workspace, a run tests your change against the commit you were
delivered. When you submit, the poster-side run puts your change on top of the
repository's default branch as that branch stands then, on a branch of its own
for the poster to review. If the default branch has moved since delivery, the two
results can differ. A run inside a claim workspace prints a line reminding you of
this, unless you pass --json.
Options
| Flag | What it does |
|---|---|
--claim <id> |
The claim this work belongs to. Inside a claim workspace, read from the claim ledger |
--target <git-url> |
Repository the work is verified against. A whole-repository delivery supplies its own |
--sha <40-hex> |
The commit your diff applies on top of. Inside a claim workspace, the commit it was delivered from |
--slice <a,b,c> |
Optional local guard rail: files you mean to touch. A diff touching anything else is refused before any container starts, and nothing re-derives that list at submission |
--local <dir> |
Checkout to read the working diff from (default: the current directory) |
--watch |
Re-run whenever a file in the checkout changes |
--keep <seconds> |
Hold the preview open this long (default: until you press Ctrl-C) |
--no-preview |
Skip the preview URL |
--no-screenshots |
Skip the before/after screenshots. See Screenshots |
--preview-route <r> |
Routes to screenshot, comma-separated (default: /; at most 5) |
--json |
Print the result as JSON instead of a report |
--test-command <cmd> |
Override the derived test command. The override is recorded in the result |
--placement <kind> |
Where the run happens: local-docker (default) or hosted. See Where it runs |
How a run is fenced
A run is two steps, and they are fenced differently on purpose.
Install gets a network, through one door. Dependencies have to come from somewhere, so the install step reaches the outside only through an allowlisting proxy on its own internal network. There is no route around it.
A lockfile that falls behind is installed around, and named. If a claim adds
or changes a dependency without including the lockfile, package.json asks for
entries the lockfile does not list. A frozen install
(npm ci, yarn install --frozen-lockfile, pnpm install --frozen-lockfile)
would refuse that outright. When the run finds that gap, it installs without the
frozen check and prints which entries are missing. The run can still pass; the
line says to include the updated lockfile in the patch, or regenerate it before
merging. A lockfile that
lists everything stays frozen.
Tests get no network at all. The test step runs with networking switched off entirely — no interface, no proxy, nothing to reach. A test suite that quietly depends on a live service fails here, which is the point: so does the poster's CI.
A posting can set its own test time limit. When the posting you claimed has
one, measured by running its tests on Terminal Hire's machines, th run stops the
test step at that limit (5 to 120 minutes) instead of the default. The report
names it, as in time limit 45 min, with the install and test steps' own times.
The container drops all capabilities, forbids privilege escalation, and caps process count. The repository's own code never runs anywhere else.
The preview
Unless you pass --no-preview, a local run ends with a URL. Opening it shows
the same result the terminal printed — the verdict, the test output tail, and
what was actually run.
The preview binds to loopback only. It is reachable from your machine and
nowhere else, and it stops when you press Ctrl-C or when --keep runs out.
Expiry and revocation beyond that are not built yet.
The preview never changes the verdict. If it cannot start, the run says why
on a preview not started: line and returns its result as if you had passed
--no-preview. A hosted run always takes that path for now: serving a preview
from the venue would need ingress to it, which is not built.
Screenshots
When the repository can be built or served, a local run ends with before and
after screenshots. Once the tests have a verdict, the run builds your version of
the tree with no network, takes one screenshot of each route at desktop size in
light mode, and then does the same for the tree without your change. The PNGs
are saved under ~/.terminalhire/screenshots/, and the report lists where:
screenshots 2 in ~/.terminalhire/screenshots/2026-09-24T18-08-41-314Z-1w4zbq (routes /; light; desktop)
after: 1 shot
before: 1 shot
rendered with no network: no backend, no signed-in state
Which repositories qualify is read from what the project declares, not from a
list of frameworks: a build script, a preview, start, serve or dev
script, or an index.html. If the build writes an index.html, that output is
served as static files. Otherwise the run starts the serve script, again with no
network. A repository that qualifies on neither count is skipped, and the report
says why in one line, for example screenshots skipped — nothing here can be served. A build that needs the network, such as one that downloads fonts, is
skipped the same way. So is a workspace root.
The pictures show what the app renders on its own: nothing it would fetch from a backend, and no signed-in state.
Screenshots never change the test result, and they stay on your machine. The
first run builds the screenshot image, about 900 MB. --watch never takes
screenshots. --json output carries them as a screenshots object with each
file's path and SHA-256.
Your own screenshots
On a first-party posting you can add pictures you took yourself to the
submission. Take them with Playwright, or reuse the ones the run saved, list
them in a small JSON file, and pass it to terminalhire claim submit <id> --screenshots <file>:
{
"v": 1,
"shots": [
{ "file": "./shots/home-dark.png", "caption": "Home in dark mode" },
{ "file": "./shots/settings.png", "caption": "Settings, saved state" }
]
}
Paths resolve against the JSON file's directory. Up to six pictures, each a PNG of at most 4 MiB (4,194,304 bytes) and at most 8192 pixels on a side, with a caption of up to 140 characters. The CLI checks all of that before it asks you to confirm, and a bad file is refused with nothing sent. The pictures upload after the patch is applied. One that does not arrive never fails the submission; the claim page lists it as not arrived.
The poster sees them on the claim page beside the verification run's own pictures, under a heading that says they came from you. Send pictures of this patch as it runs, never edited ones.
Where it runs
By default a run happens on your machine, in your Docker daemon. That is
what --placement local-docker means.
--placement hosted asks for a hosted pool: a fresh cloud VM booted for
your run and torn down after it — teardown is attempted per run, and the
platform deletes the instance at its one-hour lifetime either way — so the
verification happens on a machine neither you nor the poster operates. It
needs cloud credentials on the machine that launches it, so it is meant for
our own dispatch worker rather than a developer's terminal — a bare th run --placement hosted without those credentials refuses before any machine is
booted, and nothing about your work leaves your machine when it does. It
refuses the same way on Windows, whatever credentials are present: the hosted
machine is reached through a forwarded Unix socket, which OpenSSH on Windows
cannot forward. A
run's record says what it can prove: a hosted record our intake accepted
carries a Google-signed machine identity the intake checked, a run on the
worker's own machine carries only what the venue reported about itself, and
an absent identity proves nothing about where a run happened — a hosted run
that failed partway, or whose result never reached the intake, leaves none.
A refusal exits 2, not 1. In this command 1 means your tests ran and failed; 2 means we declined to do something, such as run them at all. A preview that cannot start is never a refusal: the run still returns its verdict. A script can tell the two apart, and nothing we refuse to do will ever be reported as a fault in your work.
What a result is, and is not
A run tells you what the project's own tests did, on a clean tree, with your diff applied. That is useful and it is checkable.
It is not an attestation that anything ran on hardware anyone can vouch for. The result records this about itself in plain terms rather than implying otherwise. A hosted run dispatched by our worker can go one step further — when its record lands, it carries the venue VM's Google-signed instance identity, checked at intake — but that names the machine, not the code it ran, and the result keeps saying which claim it is making.
Where this fits
th run is the step before claim submit. Run it
until it is green, then submit. On a first-party posting you can read the poster's
own verdict afterwards with claim runs <id>.
If the run on your submission comes back red
Send another one. A fix is a new commit, and a new commit gets its own run — you are not limited to one submission, and you never were.
Every run on a claim is kept, and the page lists them newest first — up to the twenty most recent — each against the commit it tested. A run is numbered from your first, so a capped list never tells you the eleventh try was the first. The poster's own page lists the same runs. Three reds above a green is not a mark against the work; it is how the work got there.
A run that was refused is different from one that failed. Refused means we never got as far as testing your code, so it says nothing about it — you can send a newer commit, but there may be nothing to fix.