Skip to content

How do you give an AI agent a login without giving it the password?

Unity Horizon6 min read

Put the credential below the model. The executor fills the field after the model call, so the password never reaches the context window or the logs.

You do not hand it over. The model emits a placeholder, and the code driving the browser swaps in the real value after the model call, so the password reaches the form field and never the context window.

The mechanism is that a credential has two jobs and they belong in different places. Deciding to log in is a model job. Typing the string is not. Every design that holds puts a substitution boundary between those two, and every design that leaks is one where the model held the string.

Why is putting the password in the prompt the documented answer and the wrong one?

Because the same document tells you not to. Anthropic's computer use tool page carries both instructions under two different headings. The security warning lists "avoiding giving the model access to sensitive data, such as account login information, to prevent information theft." The prompting section says that "if you need the model to log in, provide it with the username and password in your prompt inside XML tags such as <robot_credentials>", and adds, in the same breath, that "using computer use within applications that require login increases the risk of bad outcomes as a result of prompt injection."

Then count the copies. The Messages API is stateless, so every turn resends the whole message list, and a forty-turn run carrying the credential in the first user message transmits it forty times. Mark that prefix as cacheable and Anthropic stores it for five minutes by default, or an hour at 2x the base input token price for the write. The same string then lands in your application log, your error tracker, whatever your CI pipeline archived, and the trace you turned on last Tuesday to debug something else.

A vault decides who may fetch the secret. It decides nothing about where the secret goes next.

What does the substitution boundary look like in code?

It is a dictionary and a filter. The browser-use sensitive data template maps placeholder names to real values, and the documentation states the contract plainly: "The LLM only sees placeholders (x_user, x_pass), we filter your sensitive data from the input text", while the values are "injected directly into form fields after the LLM call." Credentials are scoped per domain with a regex key, so a placeholder offered on one host is not available on another.

Two settings on that page matter more than the substitution itself. It tells you to set use_vision=False, because a screenshot is an image and a text filter cannot read one. Then it goes further and prefers storage_state='./auth.json', a saved session, over supplying a password at all. Best credential handling is no credential.

Anthropic ships the same shape in its consumer product and says so. The Chrome extension safety page lists "inputting sensitive data" among the things Claude will not do, and notes the exception: "with 1Password for Claude, Claude can complete tasks that require signing in without handling the credential itself." Otherwise the extension rides your existing browser logins.

Two products and one API reference. The API reference is the outlier.

Does the substitution boundary actually hold?

Not on its own, and there is an open bug that shows why. browser-use issue 5592, filed on 29 August 2026 and still open when we checked on 29 September 2026, reports that collect_sensitive_data_values() flattens domain-scoped credentials into a single dictionary. Two domains using the same placeholder name overwrite each other, and in the reporter's words "the dropped value is never redacted from LLM-visible text." The replacement path is domain aware. The redaction path is not, so at least one secret survives in plain text across message history, extracted content and logs.

Read that as the general case rather than one project's defect. Substitution needs two mirrors kept in sync: the one that puts the value into the page, and the one that strips it out of everything the model and your logging stack can see. They are written months apart, usually by different people, and only the first one has an obvious test.

Redaction looks solid until two sites both call the field password.

What are the four options, and what does each one still expose?

Approach What the model sees Runs unattended What still leaks
Credential in the prompt the password yes message history, prompt cache, logs, any trace
Placeholder substituted by the executor x_pass yes anything the redaction path misses, per issue 5592
Human takeover at the login step nothing no nothing, until the retained session is stolen
Saved session state, no password anywhere nothing yes the stored cookie file, and its backups

Takeover is OpenAI's answer in Operator and in the ChatGPT agent. OpenAI's own documentation describes it the same way in both: the agent asks the person to take over when sensitive information has to be entered, such as login credentials or payment details, and nothing the person enters during that window is collected or screenshotted. The agent keeps the session afterwards and carries on. Cleanest boundary on the list, because no machine context ever holds the secret. It also ends unattended operation, and a filing that runs at three in the morning has nobody at the keyboard to hand the browser to.

Who does this cost, and who never notices it?

The target system decides, not the agent framework. A team integrating with a product that issues API tokens or OAuth grants has no password to protect and can skip this article. A team driving a public portal whose entire authentication story is a username box and a password box pays for a substitution layer, a redaction path, a session store, a rotation procedure and someone who remembers to run it the week an employee leaves. On every integration, for as long as that portal exists.

We built OFFSET, which files declarations directly to the ΑΑΔΕ portal. OFFSET encrypts credentials and guest data end to end. The browser-driving pattern underneath it is the one described here, and the portal has never been asked whether it would prefer a token.

What is left to worry about once the password is out of the context window?

The page. Anthropic published browser injection measurements on 26 August 2026: attacks that reached the model succeeded against Claude Opus 4.5 17.6% of the time and against Claude Opus 5 3.8% of the time, before additional safeguards. With probes and the safety classifier running, no attacks succeeded against Claude Sonnet 5, Claude Opus 5 or Claude Mythos 5, and 0.3% succeeded against Claude Fable 5. Its earlier research post is blunt about how to read a small number: "a 1% attack success rate, while a significant improvement, still represents meaningful risk", measured against an adaptive attacker given 100 attempts per environment. Figures checked by Unity Horizon on 29 September 2026.

A rate per encounter is not a rate per run. An agent that visits forty pages meets those dice forty times, and an agent that runs twice a day meets them again tomorrow. Which is the whole argument for the substitution boundary. It leaves the injection rate exactly where it found it and changes the size of the prize: the most a successful injection can read out of the context window is a placeholder scoped to a host it is not on.

We have not red-teamed our own portal agents, and the rates above are Anthropic's, measured on the open web rather than on a logged-in government form. A portal behind authentication is a narrower attack surface than an arbitrary web page, and we do not have a number for how much narrower. If someone has measured that, we would rather publish their figure than reason around it.

So run the search. Grep your stored agent transcripts for your own password, and if it comes back with a hit, drawing that boundary is work we scope. Which of the four rows above is your integration missing today?

Common questions

Can you give an AI agent a password safely?

Not by putting it in the prompt. The safe shape is a placeholder the model sees and a real value the executor substitutes into the form field after the model call, so the credential never enters the message history, the prompt cache, or anything downstream that reads either of them.

Does Anthropic recommend giving Claude login credentials for computer use?

Its computer use documentation says both things on one page. The security warning advises avoiding giving the model access to account login information. The prompting section says to supply the username and password inside robot_credentials XML tags if the model needs to log in, and warns in the same paragraph that logging in raises the risk from prompt injection.

What is takeover mode in a browser agent?

The agent pauses at the login step and hands the virtual browser to the person, who types the credential themselves. Nothing entered during that window is screenshotted. The agent keeps the session afterwards and continues. It is the strongest boundary available and it rules out unattended scheduled runs.

Is a secrets vault enough to secure an AI agent's credentials?

No. A vault controls who may fetch a secret and audits the fetch. It says nothing about where the secret goes after retrieval, which for an agent is the model's context window, the transcript resent on every turn, the prompt cache and the run logs.

Ready when you are.

Tell us what slows your business down. We'll show you what intelligent software can do about it, usually within a week.

Book a callgiannis@unity-horizon.com