Hark Handoff Is a Cheap Computer Use Agent With Big Claims

Brett Adcock's Hark unveiled Handoff, a computer use agent that claims the top-ever Online-Mind2Web score at a tenth of the frontier token price. The benchmark comparison has caveats worth reading before you buy in.

Hark Handoff Is a Cheap Computer Use Agent With Big Claims

Hark Handoff is a computer use agent from Brett Adcock’s new AI lab, and it’s making two big claims. First, the top-ever score on a web agent benchmark. Second, a token price a tenth of the frontier. Both deserve a closer look.

Hark announced Handoff on August 5 with its first research preview. Sign-ups opened the same day, with availability planned for later this summer.

What Hark Handoff does

Handoff is a browser agent built for long, autonomous web tasks. Order dinner on DoorDash, book a flight on United, message job candidates on LinkedIn. Hark says each request spins up a dedicated virtual computer with its own browser, file system, and terminal.

You connect existing accounts so the agent logs in with your saved addresses, payment methods, and history. It navigates sites with no official APIs, reading page structure and visual data instead of waiting for a documented interface.

That’s the part that matters for real use. Most browser agents break on sites nobody built for automation. Hark’s pitch is that Handoff handles the open-ended web, not a fixed set of demo sites.

The benchmark asterisks

Hark claims Handoff scored 97.7 on Online-Mind2Web, a human-evaluated web agent benchmark. That’s against 92.8 for OpenAI’s GPT 5.4, 84.1 for Claude Opus 4.8, and 69 for Gemini 2.5 Pro, according to Hark.

Here’s the asterisk. VentureBeat flagged that those comparisons are against the previous generation of frontier models. GPT 5.5, GPT 5.6, and Opus 5 haven’t published Online-Mind2Web results.

That means “top ever” can’t be checked against the strongest current systems. The newest frontier models posted their biggest gains precisely in computer use, so the gap may have narrowed or closed since Hark’s runs.

Even inside Hark’s own benchmark table, the story is mixed. On WebTailBench v2, one of the three benchmarks Hark lists, GPT 5.5 scores 72.3 to Handoff’s 68.6, per VentureBeat.

The price is the real story

Handoff runs at $0.18 per million input tokens and $2.37 per million output tokens. GPT 5.5 lists at $5 and $30. That’s a tenfold price difference, and Hark claims 0.8 second per-turn model latency.

The pricing advantage holds even against newer models. Anthropic’s Opus 5 carries the same $5 and $25 list prices as its predecessor, so Handoff’s cost edge survives the frontier comparison even if the benchmark claim doesn’t.

Hark admits it has only done post-training so far. Pre-training is “planned for later this year,” and the company hasn’t said which base model Handoff runs on.

The open questions

Security is the big one. Who can access the dedicated virtual computers and the files created on them? Hark says security and privacy are a primary focus, but the preview doesn’t specify access controls, retention, or who can inspect a session.

That matters because Handoff connects to your real accounts. A browser agent with your saved payment methods is a different risk surface than a chatbot that just reads your prompts.

Hark says 74.9% of all computer use happens in a browser, per its research preview. That stat explains why the company leads with a browser agent instead of the hardware it has hinted at. The web is where the work is, so that’s where the agent starts.

Handoff is Adcock’s fourth company. He sold Vettery, built Archer Aviation, and founded Figure AI before raising $700 million for Hark at a $6 billion valuation, TechCrunch reported.

That war chest is why Hark Handoff exists as a product so early. The company has been in stealth for months and signaled hardware ambitions, but a browser agent reaches users before any device ships. Software first, hardware later.

The computer use agent space is heating up. ChatGPT Work already ships multi-hour agent sessions, and OpenAI’s Sol ran real businesses in testing.

Hark Handoff is entering a crowded field with a price advantage and a benchmark gap to prove.

My take

The price is the headline. If Hark Handoff performs anywhere near its claims at a tenth of the cost, it changes the math for anyone building on top of agents. That’s the part I’d bet on being real.

The benchmark claim needs a lot of salt. Comparing against last generation’s models and calling it “top ever” is marketing, not measurement. The honest read is that Hark Handoff is a serious contender on price and latency, with an unproven ceiling.

The pipeline is worth watching too. Hark says it used supervised fine-tuning followed by reinforcement learning with the GRPO algorithm, per VentureBeat.

Post-training only, no pre-training yet, and no disclosed base model. That’s a team iterating fast on someone else’s foundation, which is fine for a preview and risky for a product you hand real accounts to.

For builders, the move is to get on the waitlist and test it yourself. For everyone else, wait for the end-of-summer release and the third-party benchmark runs. Cheap is only cheap if the agent finishes the job without wrecking your accounts.

Tony Simons

Reviewed & Written By

Tony Simons

Independent tech reviewer and creator of Tony Reviews Things. 14 years of hands-on testing, software auditing, and workflow automation. I test the gear so you don't waste your money on junk.

Submit a Take

Your email address will not be published. Required fields are marked *