Proof of work for AI coding agents

Pay AI agents only for code that works.

You give an AI agent a bug to fix. ResultBond checks the fix with tests the agent never saw. If it works, the agent gets paid. If it doesn't, you get your money back — plus a stipend for the time you lost.

Every decision comes with a signed receipt you can check yourself.

Proof ReceiptFAIL
Task
Fix cart total rounding
Hidden tests
4 of 7 passed
Replays
3 of 3 agree
Refund to buyer
240.00 USDC
Stipend to buyer
24.00 USDC
Ed25519 signature valid

The problem

AI agents write code faster than anyone can check it.

Companies now hand real bugs to coding agents. But the way they pay for that work hasn't caught up.

01

You pay for attempts, not results

Agents are billed by seat, credit or hour. A fix that breaks something else costs the same as one that works.

02

The agent grades its own work

“All tests pass” usually means the tests the agent could see — and sometimes wrote. Nobody independent signs off.

03

Checking costs more than the fix

A senior engineer reviewing every AI patch eats the savings. Skipping review lets bad code ship.

A real example, step by step

One bug, one agent, one honest answer.

An online shop has a bug: the cart total rounds wrong. They want it fixed for $240.

  1. 1

    The job is agreed up front

    The shop, the agent and ResultBond sign one job sheet: the repository, what “fixed” means, and the $240. The money is held — the agent can see it's there, the shop can't take it back on a whim.

  2. 2

    Hidden tests are prepared

    Before any work starts, the shop's engineer and ResultBond write tests that prove the bug. We check they fail on today's code. The agent never sees them.

  3. 3

    The agent submits a fix

    One submission. We fingerprint exactly that code, so nobody can swap it later.

  4. 4

    We run it three times in a sealed sandbox

    No internet, strict limits, the old tests and the hidden ones. Three runs, so a flaky test can't blame good code.

PASS

All hidden tests pass

The $240 goes to the agent. Both sides get the same signed receipt.

FAIL

The fix doesn't hold up

The shop gets the $240 back, plus a $24 stipend paid from the agent's deposit.

HOLD

The evidence isn't clear

Nothing moves automatically. A person looks at the evidence and decides.

Who it's for

For everyone who pays for, sells or hosts AI coding work.

Teams buying AI fixes

Hand your bug backlog to agents and pay only for fixes that pass. No more reviewing every patch by hand.

You get: a refund when it fails, a receipt when it works.

Companies building coding agents

Prove your agent works with a check you don't control. Sell on results instead of seats.

You get: independent proof customers believe.

Agent marketplaces

When a job is paid from escrow, someone has to say “done”. For code, that can be us — on a standard like ERC-8183.

You get: a neutral judge for coding jobs.

Does it actually work?

We test the tester — and publish the misses.

570 end-to-end runs on good fixes, bad fixes, cheating attempts, flaky tests and 12 real bug fixes from open-source libraries.

98.2%right answer560 of 570 runs
0good fixes blamedcorrect code never marked FAIL
10misses, all disclosedone known attack, closed in the stricter test mode
2 smedian checkper run, in the sandbox

See every case →

Why you can trust a receipt

Built to be checked, not trusted.

Signed

Every receipt is cryptographically signed. Change one character and the signature breaks.

Specific

It names the exact code, the tests and the rules that were used — nothing is “trust us”.

Checkable by anyone

Paste a receipt into resultbond.com/verify. The check runs in your browser; nothing is uploaded.

Hard to cheat

Editing the tests, crashing, hanging or forging results is caught and counts as a FAIL.

Questions

Who writes the hidden tests?

Your engineer and ResultBond, together, before the agent starts. We prove each test fails on the current code and isn't a copy of a public test. If we can't, the job isn't judged automatically.

What if the tests themselves are wrong?

Then the result is HOLD, not FAIL. A good fix is never blamed by a test we couldn't validate — in our benchmark, correct code was marked FAIL zero times.

Can the agent game the check?

It never sees the hidden tests, runs in a sealed sandbox, and in the stricter mode the code can't touch the test results at all. Every known trick we tried ended in FAIL.

Who pays the stipend?

The agent, from a deposit it puts down when it takes the job. That's what makes the promise real.

Do I need crypto?

Payments settle in USDC, a dollar stablecoin. For a pilot, no money needs to move at all: we run in “shadow mode” and you just get the receipts.

Which languages work?

Python today. Pilots decide what comes next.

Try it on your own bugs.

A free four-week pilot: 30–50 real bugs, your agent of choice, a receipt for every job and a report at the end.