How I built HireZero, an AI marketing employee, for the OpenClaw 2.0 hackathon
My AI Worth Using × OpenClaw 2.0 entry is an AI marketing employee that asks before anything goes out. Here is how it’s built, the five methods that kept a fast build honest, and what still doesn’t work.

HireZero is my entry for the AI Worth Using × OpenClaw 2.0 hackathon, and the brief was Build Your Startup’s First Hire. I picked marketing, because it’s the first job a small team drops. What I built is an AI marketing employee that works shifts, brings back drafts with the sources behind them, and doesn’t post, send or spend anything until a person approves it.
The repository’s first commit was on September 9, and submissions closed on September 28. By then it had close to 700 commits. You can watch the 68-second demo, read the code on GitHub, or get HireZero free while it’s free for a limited time. This post covers how it’s built, and the methods that kept a fast build honest.
The brief, and the rules I built against
The hackathon asked for a real hire doing real startup work: built on OpenClaw 2.0 with native multiplayer, with public MIT-licensed code, with usage reported through the official Agent Index client, and with a demo of at least a minute.
One rule shaped the design more than the others: don’t generate artificial usage. Every token on the leaderboard had to come from real work. That put metering and receipts at the centre of everything.
What HireZero is made of
- The cockpit is a React and TypeScript workspace. You chat with the employee there, review its work and make decisions, on your desk or your phone.
- The host is a .NET service. It owns the ledger (SQLite plus Markdown), runs shifts, enforces permissions and meters every model turn.
- The employee is an OpenClaw agent running on Plow’s OpenClaw base image. Plow provides identity, messaging and the model route.
It ships as one Linux container image. Run it hosted, or run the open-source code on your own machine.

Method 1: the model gets no tools
The most important decision was taking the tools away. A shift is a loop (sense, prioritize, create, align, launch, measure, decide, learn), and the host runs that loop, not the model.
Plain code does the sensing, the alignment checks, launch QA, measurement and decisions. The model does only what needs judgment: choosing priorities, writing, critiquing its own work, and the end-of-shift report.
Each model stage is one bounded turn that must return strict JSON. Permissions live in code, so the model has nothing to escalate. Before you see a draft, a critique turn scores it from 1 to 5 against a rubric, and anything scoring 3 or lower gets revised.
Method 2: approval is a boundary, not a button
Approving never publishes. That line is in the design documents, and the product follows it.
- Approve records your decision. Approve and schedule, save as draft and copy to the network’s composer are separate controls.
- Only the exact text you approved, identified by its digest, can go out, and only once.
- Uncertain outcomes wait for a person. If it can’t tell whether a post went out, it doesn’t guess.
- A chat message saying “looks good” isn’t an approval.
This makes it slower than a tool that posts for you, and that’s the point. You keep the final say.
Method 3: receipts before claims
Tokens are reserved before a turn is dispatched. An uncertain turn holds a 25,000-token reservation until it’s reconciled. Missing usage never counts as zero, and inconsistent receipts fail closed. Chat, campaigns and shifts all go through one execution gate, so nothing can spend around the meter.
That paid off on the leaderboard. The listing shows 113K tokens, rounded from 112,550 tokens in actual provider receipts over two days. That’s one owner’s real work, not traffic, and every token of it can be traced to a receipt.
Method 4: build with AI agents, verify like a skeptic
I built HireZero with AI coding agents: Claude Code and OpenAI’s Codex, often at the same time on separate parts, with a written split of who owns which files. 262 of the commits carry a Claude co-author line.
Agents make it easy to go fast. These rules kept that speed from turning into confident nonsense:
- Routine checks never call a live model. Scripted fixtures stand in for the model: no GPU inference, no paid CI minutes. A local check command records source hashes and results in a fresh receipt folder. The backend suite reached 1,282 passing tests.
- Reproduce first, then fix. A fix counts only once the failure has been seen, and then seen to be gone. One hosted entrance bug returned a 403 before its fix and a saved 200 after it, on the packaged image, not just in a unit test.
- Budget the machine. Builds check free disk first and refuse to run below a 10 GiB floor. VM tests reuse pinned images with small overlays. Every run cleans up its disposable files and keeps compact logs and hashes.
- Write down what isn’t proven. Handoff notes separate what was verified from what is only believed. That habit wrote most of the section below.
Method 5: the promo video is code
The 68-second promo isn’t a screen recording. It’s an HTML timeline on a 120 BPM, 34-bar clock, so every cut lands on a beat. Chromium seeks each frame deterministically and pipes it straight to ffmpeg: 1080p at 30 fps, in about five minutes on a CPU.
The music started as an offline synth. The final bed came from the YuE2 music model, driven by a hand-written ABC score and cleaned up with Demucs. The video mixes real UI captures with motion recreations of sample content, and labels them as such.
What went wrong
- The first real hosted shift failed at once. The model route didn’t support the “low” thinking level the configuration asked for. Turning thinking off fixed it.
- The self-review rubric can be wrong. It flagged a 92-word post as missing an 80–120-word requirement. That defect is known and still open.
- A long review hit the output cap. A queued website-copy review ran into the 4,096-token limit.
- Two things aren’t proven yet. One-click installs for outside users are waiting on Plow’s admission. Native multiplayer with two real people hasn’t been verified end to end. There are team roles, but I won’t call multiplayer done until two people have used it together.
- The writing is competent, not memorable, yet. The employee is useful. It isn’t a great copywriter.
What’s next
HireZero is brand new, free for a limited time, and open source under the MIT licence. What it does next depends on the people who use it. Send feedback, and I’ll read it. You can also open an issue on GitHub, or get HireZero free and give it a shift.
I’m Mark Hall, and I build AI tools that ask before they act. There’s more about me on the developer page, and more of my writing on markbhall.dev.
HireZero is free for a limited time. A marketing employee that brings you drafts and asks before anything goes out. Tell us what you think.
Get HireZero free →