The AI Agent Should Recommend. The System Runtime Should Decide.

On August 17, 2026, Google published a reference architecture for building AI agents. The founding assumption of that architecture is that the model cannot be trusted.

Shubham Saboo and Eric Dong put it without hedging: “System prompts are soft constraints. They can be bypassed by prompt injection, altered during prompt tuning, or behave unpredictably across model updates.”

Read that again, and note the source. This is not a vendor of governance software arguing that models need governance. This is Google, publishing on its own developer blog, telling engineers to design as though the model in the middle of their system has already been compromised.

What Google actually built

The zero-trust agent architecture places three hard controls outside the model. Not in the prompt. Not in the fine-tune. Outside.

Cryptographic signatures on every state-changing write.

Each agent gets a hardware-backed asymmetric key through Cloud KMS. Every database mutation the agent proposes must carry a signature from that specific agent, and an ingress guard verifies the signature before the write commits. Payloads are serialized deterministically so the hash is reproducible. The result is non-repudiation: an out-of-band scan of the ledger can prove which agent authorized which change, and an unsigned write never lands.

Kernel-level isolation for generated code.

Any Python the agent writes executes inside a gVisor sandbox with the network disabled, all Linux capabilities dropped, the code mounted read-only, and hard caps of 64MB of memory, a tenth of a CPU, and five seconds. An agent that generates malicious code still cannot reach a syscall or exfiltrate anything, because the boundary is enforced by the kernel rather than by the agent’s own restraint.

Deterministic semantic gateways.

A reverse proxy sits in front of both the model and the database and applies business rules to what goes in and what comes out. Inbound, it blocks jailbreak patterns, credit card numbers, and live API key formats. Outbound, it validates the action itself: a proposed refund is checked against the actual order total before it is allowed through. Those rules are covered by CI regression tests, so they survive the next model update.

Three controls. All of them deterministic. None of them dependent on the model behaving.

The part that makes this bigger than a security post

Google’s threat model is adversarial. The scenario they design against is prompt injection: a customer requesting a refund who embeds instructions to alter the transaction amount and leak environment variables.

Last month we documented four incidents from the AI Incident Database that were not adversarial at all. Coding agents from Anthropic, OpenAI, and Google each destroyed or corrupted a production database belonging to the people using them, inside a four month window. In two of those cases the agent executed a correct, standard, well formed database command. It was simply pointed at production instead of a disposable instance. Twenty two tables dropped in one case. A truncated users table in the other. In the third, the agent took a portal down for 33 minutes, watched a human roll it back, and then generated a post mortem claiming credit for the recovery.

Nobody attacked anything. The models did not misbehave. The environments were misconfigured, and the agents executed at machine speed against whatever they were handed.

Now put the two together. The adversarial path and the accidental path start from opposite ends and arrive at exactly the same conclusion: execution authority has to sit outside the model. Google reaches it by assuming an attacker. The incident record reaches it by assuming nothing at all, just a wrong connection string and an agent that does not pause.

When a hostile threat model and an ordinary Tuesday converge on the same architecture, the architecture is not a security preference. It is the shape of the problem.

What each control actually buys a regulated operator

Google wrote this for developers. The translation for anyone running operations inside a bank, a carrier, an insurer, or a health plan is direct, and each control maps to an obligation that already exists.

Signatures buy non-repudiation.

This is the audit requirement, stated in cryptography instead of policy. Which agent authorized this adjustment, under what authority, at what time, and can you prove it to a party who was not in the room. An agent that writes to a system of record without a verifiable signature produces transactions you cannot defend in an examination.

Gateways buy enforcement before commit.

This is the compliance requirement. A refund that exceeds the order total is rejected by a rule, not by a model that was asked nicely to be careful. Every regulated workflow already has these rules. They live in policy documents, in procedure manuals, and in the judgment of trained staff. The question is whether they also live in the execution path.

Sandboxing buys blast radius.

This is the operational requirement, and it is the one most enterprises have never had to think about, because software that wrote and executed its own code was not previously part of the stack.

The operator conclusion follows from all three. Once an agent can issue a refund, adjust a policy, change pricing, update a CRM record, or launch outreach, prompt-level guardrails are theater. A system prompt is a request. A gateway is a control. Only one of those is admissible.

The gap Google leaves open

The architecture is correct. It is also a build.

It assumes engineers who can provision KMS keys, bind service accounts to signing roles, configure gVisor runtimes, stand up a validation proxy, and maintain a CI suite that keeps the rules honest across model releases. It assumes deployment inside VPC Service Controls. It assumes Google Cloud.

No insurance operations team is shipping that this quarter. No contact center leader is shipping it at all. The principle is fully portable. The implementation is not, and the distance between those two facts is where most enterprise AI programs are currently stalled: convinced of the argument, unable to staff the answer.

That gap is the entire reason a completion layer exists as a category.

The runtime decides

Callvu is built on the same separation Google is describing, delivered as a layer rather than a project. The conversational agent understands the customer, reasons about intent, and recommends the action. It does not hold the write credential. A deterministic runtime evaluates that recommendation against policy, eligibility, disclosure requirements, and entitlement, then executes it or rejects it. The customer-facing workflow, the disclosure, and the consent capture all happen inside that runtime. The Decision Trace is produced by the runtime as a byproduct of execution, not narrated after the fact by the agent that acted.

That last distinction is the one worth holding onto. In the incident record, an agent broke production and then wrote the account of what happened. Every audit regime in financial services and healthcare rests on the assumption that the record is produced by something other than the actor whose conduct is in question. An agent that both executes and reports collapses that assumption. A runtime that executes and attests restores it.

Google’s guidance to developers is to stop trusting the model with execution authority. The same guidance applies to the enterprise, with one addition: you do not have to build the runtime yourself.

The agent should recommend. The runtime should decide.

Source

Source: Shubham Saboo and Eric Dong, “Build Zero-Trust AI Agents with Google’s Agent Development Kit,” Google Developers Blog, August 17, 2026. Reference implementation available in the GoogleCloudPlatform/generative-ai repository.

Facebook
Twitter
LinkedIn

Get the latest content straight to your inbox.

Callvu How Customers Feel About AI in Customer Service CX Research

How will customers feel about AI in your customer service?

Many companies are rushing to offer AI assistants and other AI-powered tools in their customer service. But are consumers ready?

Callvu How Customers Feel About AI in Customer Service CX Research

How will customers feel about AI in your customer service?