The AI Didn't Steal Your Key. It Stole Your Instruction.

When an AI agent can pay, the attack moves from the key to the instruction. Three cases from a $1.5 billion signing failure to a bot talked into a transfer, a ten-second sandbox, and one boundary rule for anyone letting software near their bitcoin.

By Dalia ·

The custody mistakes most people know about involve a person being fooled: typing a seed phrase into a fake wallet, pasting a poisoned address, approving a transfer in a hurry because a message sounded official. The person was in the loop, even when the person was wrong.

The next mistake may not involve you being fooled at all. It may involve the assistant you asked for help.

Try it before you read about it

Reading about this attack and doing it are different things, so do it first. Open the intent-theft sandbox. It uses no real keys, no real funds and invented invoices. You ask an assistant to pay a Lightning invoice, then choose what happens next. It takes about ten seconds.

If you ran it, you already have the lesson in your hands, so I'll only name what happened. The private key was never touched. The wallet did not break. The invoice the agent built was valid; it paid the wrong person, because a hidden instruction on the page rewrote what the agent thought you wanted. The failure happened before anything was signed. The instruction changed.

You also saw what a careless telling of this story gets wrong: where the loss lands depends on who was allowed to authorize the spend. Let it auto-pay, and the money is gone before you know it. Stay in the loop, and the device screen shows the real destination, which doesn't match what you intended, so you reject it. Same attack, two endings, decided by one boundary.

What you just triggered

I call it intent theft. The attacker doesn't steal the key; they corrupt the message that tells the system what the key is for. In the older mistakes, you were in the loop even when you were fooled. With an agent, you may never see the poisoned instruction at all. The agent reads it, accepts it as context, and turns it into action.

The security field has names for the two halves. Prompt injection is the malicious instruction smuggled into what the agent reads: the hidden text in the sandbox. Excessive agency is the agent holding more authority to act than it should: the auto-pay button. Intent theft is what happens when the two meet over a wallet.

This is not theoretical

Agent payments are shipping. On 12 February 2026, Lightning Labs released open-source tools that let AI agents pay over Lightning without accounts, API keys or sign-up flows, and Coinbase launched Agentic Wallets, which let an agent hold funds and send payments within spending limits. The pitch across the industry is the same: let the software pay the software.

Then the incidents arrived.

An agent talked into acting. On 4 May 2026, an AI-linked wallet on the Bankr platform, tied to the Grok bot on X, was drained of about 3 billion DRB tokens, reported at roughly $150,000 to $200,000. The attacker first sent the wallet a Bankr membership NFT that unlocked its ability to make transfers. Then they posted a message in Morse code. Grok decoded it in a reply, and Bankrbot treated the decoded text as a command and executed it. About 80% of the tokens came back after the community identified the attacker. No key was stolen. The bot only needed standing authority and an instruction.

The middle layer rewriting the instruction. In April 2026, researchers published a preprint, not yet peer reviewed, that tested 428 "LLM routers," the services that sit between users and AI models. Twenty-six were injecting malicious tool calls or stealing credentials, and in the researchers' own test one router drained a decoy wallet they had planted. A co-author separately claimed on X that a client lost $500,000 the same way; the paper does not document that case. The risk does not have to live in the wallet. It can sit in the layer where you think you are talking to one system while another edits what gets passed along.

And the one that shows this is not only an AI problem. On 21 February 2025, the exchange Bybit lost about $1.5 billion, the largest crypto theft on record, which the FBI attributed to North Korea. Its multisig cold wallet was not broken. Attackers compromised a developer's machine at Safe{Wallet}, the multisig interface Bybit used, and injected code into its web interface. Bybit's signers saw what looked like a routine transfer while they were approving a change to the wallet's contract. They were not careless. They saw one thing and signed another.

The Bankr wallet shows what happens when an agent can be talked into action. The router research shows what happens when the middle layer changes the instruction. Bybit shows what happens when the human approval layer is shown one thing and asked to approve another. In all three, the cryptography held and the decision layer failed.

Structure, not virtue

Look at what the two agent cases had in common. The keys were online and reachable. Spending authority was standing and automatic. Software could authorize a spend with no human in the room. It's tempting to read the absence of disciplined holders from this list as Bitcoiners being smarter. The difference is structural.

Disciplined custody inverts all three properties, whether that means a single hardware wallet you keep yourself or a collaborative multisig in which an accountable institution holds one of the keys. The signing keys sit offline or with separate parties. Authority is not standing; every spend is a deliberate act. And a person approves each spend on a signing device.

Bybit is the warning that inverting those three is not enough on its own. Its keys were cold, its signers were human and present, and it still lost $1.5 billion, because the signers trusted the interface instead of checking what the device was actually being asked to sign. The last defense is the one a poisoned page can't reach: the device's own screen, read by a person who knows what they meant to do. Agents didn't make that obsolete. They made it the thing standing between a convenient assistant and a six-figure lesson.

The boundary rule: assist, never authorize

If you're going to let an AI agent near your bitcoin, and many people will, because the tools are useful, the whole game fits in one line:

An agent may help you build a transaction. Only people, approving on their own signing devices, should be able to sign one.

Everything below is that rule applied.

The seed-phrase rule, extended

No screen, no message and no person ever needs your seed phrase. That rule still holds. Add one line:

No agent needs your keys, your signing device, or standing authority to move your funds. An assistant that asks for any of the three has crossed from helping into custody, and your AI assistant is not a custodian.

Like the best custody rules, it isn't a workflow you maintain. It's a line you don't cross.

Where this meets inheritance

Custody plans and inheritance plans already separate owning bitcoin from being able to access it. An agent with standing spend authority is a strange third category: not an owner, not an heir, but something that can move funds while answering to whatever it last read. When you design a custody or inheritance plan, write the boundary in: what software may see, and what only a person may authorize. Otherwise the convenience you add today becomes the hole someone else's instruction walks through later.

Which piece of software in your Bitcoin life can already act without asking you first?

An agent can carry your message. Deciding which message gets signed stays with a person and a device the page can't reach.

Sources

Disclosure: I am one of the advisers at The Bitcoin Adviser, which offers collaborative multisig custody. This essay does not recommend any provider.

Created by Dalia · More essays · bitcoinsovereign.academy