How submission verification works
When a developer submits work on a Terminalhire posting, we run the submitted commit in a fresh Google Confidential Space virtual machine. The poster can inspect the result and any checked evidence of its run environment before deciding whether to accept the work.
Confidential Space is the production worker’s configured default for new hosted submissions. Each run’s own evidence tells you what actually happened: earlier runs and checks started locally can use a different environment.
From patch to evidence
Developer submits a patch
↓
A commit on a separate branch is queued
↓
A fresh Confidential Space VM boots for that run
↓
Dependencies install through an allowed list of hosts
↓
External network access is cut; the measured tests run
↓
The intake checks the run’s evidence and records the result
↓
The poster reviews the change and decides whether to accept
The worker checks out the commit named by the dispatch. The result is bound to that commit and claim. Our intake rejects a result for the wrong commit, a superseded dispatch, or evidence replayed from another run.
The VM is torn down after the run. Cleanup failures are reported, and a VM lifetime limit provides a further deletion backstop. A failed dispatch or an unposted result can leave no recorded verdict; absence is not a passing result.
What Google Confidential Space adds
An ordinary hosted run can present a Google-signed identity for its VM. Confidential Space adds signed claims about the venue workload image and the hardware and boot environment. Our intake checks:
- Google’s signature, the intended audience and the run’s one-use nonce;
- the expected project, zone and instance;
- the venue workload image digest against our approved list;
- the supported confidential-computing hardware, Secure Boot and disabled debugging.
The worker also binds the attested workload to its connection before sending code to the venue. A wrong image, mismatched run or invalid token is refused. The Confidential Space path does not silently fall back to an ordinary VM.
This gives the poster evidence about the environment that produced a result. Google does not certify that the patch is correct. The approved digest identifies the outer venue workload; it is not a recorded digest of the inner container image chosen to run a repository’s tests.
Confidential Space also does not change the developer’s repository access. A whole-repository posting gives the assigned developer read access to its files, history and branches. It is not a promise that Terminalhire cannot access source.
Install first, then run the tests offline
Dependencies need to be downloaded. During installation, the container can reach only the configured package hosts through a proxy outside the container. The container cannot change that proxy’s allowed list.
The measured test phase runs with Docker’s --network=none: no external network
access. Tests that depend on a live service need an offline fixture or a different
test setup; a missing service is not quietly opened up for the suite.
For Java, dependency preparation can run the suite once through the restricted proxy in a throwaway copy. That preparation pass is not the recorded verdict. The measured run is still the offline one.
Hosted test containers also run as a non-root user, drop Linux capabilities and disallow new privileges. The staged workspace stays writable so normal build tools and tests can create files. We do not supply production application data or credentials for tests.
Where the test command comes from
A posting the poster reviews by hand has no test command and gets no run: see
Post work. Every other posting needs one. At publish time we read the
repository and look for one. When we find none, the poster names it, and picks
the language when we cannot tell. The poster cannot replace a language we found, or a
command we found that runs only the tests; when the command we found also runs checks such
as lint or typechecking beside them, the poster may name one that runs only the tests, and
verification runs that one. A command the poster names must be plain: a build or test tool,
or a ./ script, with words, flags and relative paths. Before they claim, developers see
whether the command was found or named, and they see the command itself unless it is a
found one that is not plain; that one they read in the repository after claiming. It
cannot change once someone has claimed.
Reading a result
The claim’s verification panel and the poster tools show accepted run records. Read the result alongside the diff and any separate GitHub CI checks.
A result can include the submitted commit, test command and its source, exit code, test counts, output tail, cleanup state and checked venue evidence. The panel names the venue kind when the run also returned a readable venue descriptor. Missing evidence must not be inferred from the deployment default.
The panel and claim_verification lead with that record: the test command, who
chose it (the repository, the poster, the developer or terminalhire), the commit
and the image. When the command recorded no exit code, the record says so rather
than saying it ran. The pass and fail counts come after it, labelled as our
reading of the output, which anyone with the output can recount. When the command
exited but we cannot read counts from its output, the panel says the gap is ours,
not the developer’s.
When a run did not pass and its output shows npm's or pnpm's step lines, the
panel and claim_verification also name the last package script named in the
output before the command ended, such as check:bundle. This tells you where to look. It is not a
finding: a script that runs another script names the inner one, and the output
can name the wrong step.
When our worker signed its reading of a run with a key we hold, the claim page
and claim_verification add one line: “Signed by Terminal Hire's worker: this
is our worker's own reading of the run, and it does not show what ran on the
machine.” The line appears only on runs our worker signed. It shows who wrote
the record, not what the machine did.
Test counts come from the summary a test runner prints. When a command runs more than one runner, or one runner prints a summary for each file, the output holds several summaries. The counts then come from one of them and are marked partial: the panel says they are not a total for the whole run. The result itself still follows the exit code and any failures the run reported.
| Result | What it means |
|---|---|
| Tests completed successfully | The observed suite passed for that run. Review the diff and the scope of those tests. |
| Tests failed | The suite reported failures. Read the command and output to investigate. |
| No tests observed | The run did not establish a passing suite. |
| Time limit reached | Our time limit, or the posting's own test time limit when it has one, ended the run. It says nothing about the work. |
| Environment failure, unavailable test command, or unparsed counts | Our verification path could not render a reliable verdict. |
| Refused | A required verification condition was not met. This is not a verdict that the patch is broken. |
For command-line callers, failed tests and no tests observed use exit 1. Refusals, our time limit, and infrastructure or interpretation failures use exit 2. These are distinct from successful completion.
When the poster named an acceptance command, the panel and claim_verification add one line for it: the exact command and what happened to it. If the full command fails, we run it on its own, on a fresh copy of the submitted code. The result above is still the repository's full test command. The acceptance command is the check the poster named for this task.
Passing tests do not accept a claim or release payment. The poster decides whether to accept and pay under the posting’s terms, and separately when to merge.
Checks you start on your own machine
A check started from your terminal defaults to local Docker. It uses the same install-then-offline containment, but that does not establish a Google-hosted or Confidential Space identity.
Its output stays with the person who started it; it is not automatically posted as the claim’s verification evidence. The local preview is bound to loopback, so sending its link to someone else does not share the result. A local check helps you prepare the patch; the submitted commit gets its own worker-dispatched run.
| Path | Environment | Where the result goes |
|---|---|---|
| New production submission | Fresh Confidential Space VM, as configured by the production worker | Accepted evidence is recorded against the claim. |
| Earlier hosted run | May use ordinary Container-Optimized OS and Google VM identity | Read that run’s recorded evidence. |
| Local check | Docker on the machine that starts it | Local terminal and preview. |