Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

C7 — Design least-capability tools

From cognokratos/arktos-wallet · docs/capability/07-design-least-capability-tools.md · pinned revision 92650a034799

A narrow API is a security boundary.

Question: given a cryptographic capability you want an agent to have, what is the narrowest tool that delivers it? How do you tell when a tool's real authority is wider than its name?

Prerequisite: guardrails filter what goes in and out of a model, as covered in simple-agent-template stage 5. This lesson is about the layer under the guardrails. A guardrail can be bypassed by one cleverly worded input. A capability that was never built cannot be invoked by any input.

Mental model

The authority of a tool is the set of effects reachable through its arguments, not what its name promises. Three properties make that set small:

  1. Specific operations. The tool's name and its schema fix what happens. The arguments select only which resource, within a range the server enforces.
  2. Typed, schema-described contracts. Inputs and outputs are typed, validated and schema-described. The server decides the output shape, never the caller.
  3. Public-only results. Nothing value-bearing comes back, whatever the arguments.

How Arktos tools are built

PropertyWhereEffect
Typed requestsCreateWalletRequest, GetBitcoinAddressRequest, GetEthereumAddressRequest in src/wallet_services.rsTwo fields at most: wallet_name (validated WalletName) and account_index (0..=2^31-1, enforced by DerivationIndex and declared in the schema). Request types do not use deny_unknown_fields: unknown arguments are currently ignored, not rejected, and none of them can select an identity (lesson 06)
Typed responsesCreateWalletResponse, BitcoinAddressResponse, EthereumAddressResponse, marked PUBLIC-ONLYNo secret-bearing field exists to fill. deny_unknown_fields makes deserialization into these Rust response types reject unexpected fields, which lets the tests detect response-shape drift
Generated JSON Schemaschemars derives, published by rmcp as inputSchema and outputSchemaThe schema is generated from the same types the server deserializes into, so there is no hand-written schema to drift from the code
Structured contentJson<…> returns in src/mcp.rsResults are data, not prose for the model to parse
Honest descriptions#[tool(description = …)]Side effects are stated ("Creates state", "the account is recorded on first use") along with what is not returned ("the recovery phrase is never returned")
Two error channelsAppError in src/error.rsClient errors (invalid_argument, not_found, already_exists) are tool results the model can act on. Server faults are an opaque -32603
Configuration is not an argumentChainConfig from BITCOIN_NETWORK and ETHEREUM_CHAIN_IDThe model cannot select a network
Minimal setFour toolsThe whole surface fits in one table

A spectrum of tools

None of the hypothetical tools below exists in Arktos. They are future design comparisons only.

ToolCan it reveal value-bearing secret material?Abusable by prompt injection?Authority broader than its name?Constrainable structurally?Should a model call it?Human approval?
Safer: get_bitcoin_address(wallet_name, account_index) (implemented)NoOnly to fetch the caller's own public dataNoAlready is: typed, ranged, owner-scoped, network from configYesNo
Dangerous: wallet_execute(operation, payload)Whatever any operation can do, including operations added laterYes, because the attacker picks the operationYes, by construction. Its real authority is the union of everything it dispatches to, and that set grows silentlyNot without turning it back into separate toolsNo. Split it into specific toolsCannot be decided per call, because the tool has no fixed meaning
Very dangerous: export_seed(wallet_name)Yes, all of it, permanentlyCatastrophicallyIt is total authorityNo. The output is the secretNoNo amount of approval makes model-context disclosure safe
High authority: sign_arbitrary_bytes(wallet_name, bytes)Not the key, but it is a signing oracleYes: the attacker supplies the bytesYes. "Arbitrary bytes" includes serialized transactions, login challenges for other services, and attestationsOnly by replacing "bytes" with typed, domain-separated structuresNot in this formYes, but a human cannot meaningfully approve opaque bytes either

Composition can exceed the sum of the parts

Authority must be evaluated across tools, not one tool at a time:

  • get_account_xpub (privacy loss: every address becomes linkable) plus derive_private_key(index) (one account) gives you the parent account private key, and with it every account under it. Non-hardened BIP32 derivation lets anyone holding the parent xpub and any non-hardened child private key solve for the parent private key. Arktos's last two path levels are non-hardened.
  • sign_arbitrary_bytes plus any public transaction builder is sign_transaction with no policy attached.
  • create_wallet plus a future sign_transaction that is missing a per-wallet allowlist lets an agent create a fresh wallet and then spend from wallets it was never meant to touch, if wallet selection is just a name.

The operator has capabilities the agent does not

Arktos ships an operator tool, src/bin/secret.rs (make encrypt, make decrypt, make hash). make decrypt prints a recovery phrase. That is not a contradiction of this lesson. It is the lesson applied:

  • It is not reachable over MCP or HTTP. It is a separate binary.
  • It requires MASTER_KEY in the operator's own environment, plus a ciphertext the operator extracted from the database with DATABASE_KEY.
  • It exists for ceremonies such as disaster recovery or migrating a wallet to another implementation, which are run by a human who already holds the root secrets.

Capabilities are assigned by principal: the operator can do things the agent cannot, because the operator already holds the authority those things require. The learning labs never use make decrypt.

Lab

Inspect the published contract

cargo test --test mcp_protocol_tests tools_publish_input_and_output_schemas
cargo test --test mcp_protocol_tests domain_errors_are_tool_errors_with_codes

Read the first test and list every property it asserts about the schemas. Then compare GetEthereumAddressRequest with the schema you would write by hand. Is there anything in the generated schema that you would have left out?

Exercise: prove ownership of an Ethereum address

A user asks the agent: "prove to this website that I control address 0x…". Design the tool. Compare:

export_private_key(wallet_name, account_index)      → the website verifies by deriving the address
sign_challenge(wallet_name, account_index, challenge) → the website verifies the signature

export_private_key proves ownership by transferring it. After one call, the website, the model, the transcript and every log hold the key. Reject it.

sign_challenge keeps the key inside the service. Now go further: is sign_challenge safe as written? Work through each of the following:

ConcernIf missingWhat a safe design pins down
Domain separationThe "challenge" could be a valid transaction, or a challenge for a different siteA fixed, recognisable prefix or type tag (compare EIP-191's "\x19Ethereum Signed Message:\n" prefix, or EIP-712 typed data with a domain separator), so the signature can never be valid as anything else
Message structureThe attacker chooses free text that the model will happily pass alongA typed message: audience (domain or URI), statement, address, chain ID, nonce, issued-at, expiry. EIP-4361, Sign-In with Ethereum, is one existing structure of this kind
Replay protectionOne captured signature logs in foreverA verifier-issued nonce, a short expiry, and an audience the verifier checks
Purpose restrictionThe same tool becomes a general signing oracleThe tool signs only this structure, and only with a dedicated derivation path or key if possible, never "any bytes in this shape"
Who can request itAny injected instruction can trigger a signature for any sitePossibly a human confirmation that shows the parsed structure, never raw bytes

Write the request and response types for your sign_challenge as Rust structs, in the style of src/wallet_services.rs. Then say what the type itself now makes impossible.

Do not implement it in Arktos. This is a design exercise, and Challenge 2 extends it.

Explain

Why does this lesson claim that capability design matters more than prompt design? Give a concrete prompt-injection input that no system prompt reliably stops, and show which property of the tool contains it.

Failure mode

  • A generic dispatcher tool (execute, run, call) whose authority grows with every new operation.
  • "Flexible" byte-level signing APIs.
  • Evaluating each tool's risk on its own, when the risk lives in the combination.
  • Using prompt instructions ("never export the seed unless…") as the control, when the capability should not exist.

Takeaway

Capability design is more important than prompt design when agents can act.

Previous: C6 · Next: C8 — Recovery is part of security

This chapter is maintained in cognokratos/arktos-wallet beside the code it teaches. The book shows docs/capability/07-design-least-capability-tools.md at revision 92650a0347993622cbb3e5eeac2c908069c67fd6 (branch main). View source at this revision · Report a correction.

Corrections are made upstream against the current main branch and appear here when the book's pin for this source is updated.