Minimal illustration of a small agent shape sending an action through a single glowing gate before it reaches a stack of business systems.

Key takeaways

  • I messaged around 300 companies that build AI agents and asked one question: does an action get executed without anyone looking first? Eight gave me real answers.

  • Most of those eight let their agents act without a person checking first, and when something goes wrong they find out afterward.

  • The few who do check first mostly do it because they do not trust the agent yet, not because a law requires it.

  • Larger surveys show a related picture. In a Cequence and EMA survey published in August 2026, 34% of organizations said they evaluate an agent's authorization at the moment it attempts a specific action.

  • An agent can be fully authorized and still take the wrong action. The gap is closed by checking each proposed action against policy and evidence at runtime, before it executes.

I have been sending the same message to founders and engineers at companies that build AI agents. It contains one question: does an action get executed without anyone looking?

About 300 companies later, eight have given me real answers. This post covers what those eight told me, what larger surveys say around them, and what I think it means for any team about to give an agent write access to something that matters.

How did I run this research, and how far can you trust it?

The method is simple. I messaged around 300 companies that build AI agents and asked the same question each time. Eight have given me real answers so far, and I am still asking.

Eight is a small number, and I want to be plain about what follows from that. The people who replied chose to reply, so this is not a random sample. I would not turn any of it into a percentage, and I would not claim it describes the whole market. I am also not naming any company, because the point is the pattern, not the people.

What a small set of honest answers does give you is a first look at how real teams treat the moment between an agent deciding and a system changing. That moment is where I work, so I wanted to see it from the other side.

What did the eight companies tell me?

Three things came up, and they connect to each other.

Most let their agents act without a person checking first

In this group, the default is that an agent decides and the action runs. There is no reviewer between the model's choice and the write to the outside system. If an agent decides to update a record, send a message or trigger a payment, the call goes through the same way it would if a person had clicked the button.

When something goes wrong, they find out afterward

The answers described discovery after the fact, not prevention beforehand. A problem surfaces once the action has already changed something.

Observability tools are built for exactly this situation, and they do it well. They tell you what already happened. That is a valuable job, and it is a different job from deciding whether an action should happen at all.

The ones who check first mostly do it out of distrust, not regulation

You might expect companies that put a person in the loop to point to compliance. Mostly, they did not. The most common reason was simpler: they do not trust the agent yet.

Here is my own read on that, and I want to mark it as mine, because no respondent said it in these words. A checkpoint that exists because of distrust has a shelf life. Each clean run makes the reviewer look like overhead, and the natural pressure is to remove the reviewer once the agent has earned some confidence. Confidence built on a run of good outcomes says very little about the next action, which may be the one that is wrong.

Infographic showing around 300 companies messaged, 8 real answers, and the three findings from the research.

Is this just a quirk of a small sample?

It is worth checking the pattern against larger surveys. None of them asked my exact question, so they do not confirm it. They do point in a related direction.

Checking at the moment of action is the exception. According to research from Cequence Security and Enterprise Management Associates (EMA), published August 31, 2026, only 34% of organizations evaluate an AI agent's authorization at the moment it attempts a specific action. The study surveyed 202 enterprise IT and security leaders at organizations with 1,000 or more employees that are deploying or evaluating agentic AI. The same study reports that 65% of respondents have experienced an agent take an action outside its intended scope, and 29% of respondents reported measurable business impact. Cequence sells security products, so read the findings with that in mind.

Agents stepping outside their limits is common, and response is slow. The Cloud Security Alliance (CSA) published a survey report in April 2026 titled Enterprise AI Security Starts with AI Agents. It found that 53% of organizations reported that AI agents exceed their intended permissions at least occasionally, and only 8% said it never happens. The report also found that 58% of organizations said detection and response take five hours or longer. Zenity, a vendor of AI security and governance software, commissioned the study, and CSA ran the survey online in September and November 2025 with 445 responses from IT and security professionals.

Three statistics from Cequence and EMA and the Cloud Security Alliance about AI agent authorization, permissions and response time.

These surveys ask different questions, use different samples and were run by organizations with a stake in the topic, so I would not stack them into one number. Taken together with my eight conversations, the direction is consistent: a lot of agent activity runs with checks that happen before the agent is deployed or after the action, and fewer teams check at the moment of the action itself.

Why does an authorized agent still take the wrong action?

Because permission and correctness are two different questions.

A permission system answers whether this agent is allowed to call this tool. It does that job well, and nothing here argues against it. It was never built to answer a second question: is the value the agent is about to commit still true?

The core idea on our website is that authorized does not mean correct. An agent can hold the right credentials, have the right permissions and send a perfectly formed tool call, and still take the wrong action.

Here is a case from the Salus product demo, published on usesalus.ai. An agent is booking a trip. It proposes a reservation using a $299 itinerary that was saved earlier in the conversation. The call is authorized and correctly formatted. But live availability and the current total were never verified. Salus stops the action, so the booking never reaches the airline. It tells the agent exactly what is missing: search live inventory. The agent does, and finds that the real total is now $375. The customer confirms the new price, the booking runs, and a receipt is created.

Nothing about the first attempt was unauthorized. It just was not true anymore.

What does a check at the moment of action look like?

Salus sits on the path between your agent and your backend. Your agent proposes an action, and before your backend executes it, Salus checks the action, the policy and the evidence behind it. The product page describes the decision path in five stages: the proposal, structural checks on policy, evidence, provenance, limits and state, reasoning about intent when structure alone cannot settle the action, execution control, and a receipt.

The website groups the outcome into three verbs.

  • Prevent. Salus holds an unsupported write before your backend performs it.

  • Repair. Fixable actions go back to the agent with the missing evidence and the next step, so the agent can correct itself and finish.

  • Prove. Every consequential action produces a receipt that shows what was proposed, what was checked, why the decision was made, and whether execution occurred.

The booking example above is the evidence and facts check at work, one of fourteen families of checks that also cover things like action identity, duplicate control, workflow state, policy rules, and confirmation and approval.

Integration is deliberately small. You protect the tool that performs the write and leave your model, agent loop, framework and backend as they are. This is the example shown on the Salus website:

from salus import Salus

salus = Salus()

# do_refund is your existing tool implementation.
issue_refund = salus.protect(
    "issue_refund",
    do_refund,
    side_effect=True,
    risk="high",
)

The integrations page lists the frameworks and providers Salus works with, including LangChain, LangGraph, CrewAI, AutoGen, Vapi, Retell and MCP, and a provider webhook can be routed through Salus without an SDK.

For teams that are not ready to block anything, each protected route can run in one of two modes. In shadow mode, the action continues while Salus records the decision and receipt, so you can see what it would have decided. In enforcement mode, Salus allows, revises, escalates or blocks according to the active policy. The mode is set independently for each route.

That matters for the finding at the center of this post. If a checkpoint exists only because a team does not trust the agent yet, a way to observe decisions before enforcing any of them gives that team evidence to trust, in place of a feeling.

What do the benchmark results show?

Salus publishes three benchmark results on its homepage. They are lab benchmarks, not customer results, and I want to present them that way.

  • 70.6% fewer policy failures on τ² bench.

  • 40.4% more compliant completions on τ² bench.

  • 20.5% lower agent model spend per successful task on CarBench.

These figures are quoted from the Salus website at usesalus.ai, which is where they are published. τ² bench is an evaluation framework from Sierra Research for agents that use tools and follow written business policies in customer service settings. It is open source, and its repository is listed in the sources below. Treat the numbers as a measure of how the product performs under those test conditions, and not as a prediction of what any individual company will see in production.

Where should a team start this week?

You do not need a new program to begin. Five steps are enough to learn where you stand.

  1. List the tools your agents can call that change something outside your own system. Refunds, bookings, record updates and outbound messages belong on the list. Read only searches do not.

  2. For each one, ask whether you would find out about a wrong action before or after it ran. Be honest about the answer. It is the question I have been asking.

  3. Pick one consequential route and run it in shadow mode. Read the decisions for a while before changing anything.

  4. Decide what a wrong action should get. A fixable mistake should go back to the agent with the missing step. A judgment call should go to a person. Some actions should simply stop.

  5. Ask your own team the question I asked. Does an action execute without anyone looking? The answer may be more specific than you expect.

If you want help putting the first action path behind Salus, you can get started here or read the docs.

What I am still asking

I am still collecting answers. If your team builds agents and you are willing to tell me plainly whether an action executes without anyone looking, I would like to hear it, and I will keep your company anonymous. You can reach the team at [email protected].

Frequently asked questions

What is runtime authorization for AI agents?

Runtime authorization checks each action an AI agent proposes against policy and evidence before it executes. Salus does this between the agent and the backend, and either allows the action, returns an exact fix to the agent, sends it to a person, or blocks it.

What does "authorized does not mean correct" mean?

An agent can have valid credentials and permissions and send a well formed tool call, and still take the wrong action, for example by committing a price that is no longer true. Permission tells you the agent may call the tool. It does not tell you the action is right.

Do most companies check AI agent actions before they run?

In my group of eight respondents, most did not. In a Cequence and EMA survey published August 31, 2026, 34% of organizations said they evaluate an agent's authorization at the moment it attempts a specific action. Different studies ask different questions, so treat the comparison as directional.

How many companies did you ask, and how many answered?

I messaged around 300 companies that build AI agents. Eight have given me real answers so far. The sample is small and the respondents chose to answer, so it is not a statistically representative survey.

Are the Salus benchmark numbers customer results?

No. The 70.6%, 40.4% and 20.5% figures come from lab benchmarks (τ² bench and CarBench) published on the Salus website. They are not results from customer deployments.

Do I have to replace my agent framework or model to use Salus?

No. You protect the tool that performs the write, or route a provider webhook through Salus. Your model, framework, agent loop and backend stay yours.

What is shadow mode?

In shadow mode, the action continues as normal while Salus records what it would have decided, along with a receipt. In enforcement mode, Salus acts on the decision before execution. The mode is set for each route independently.

About Salus

Salus is runtime authorization for AI agents. It checks consequential agent tool calls against policy and evidence before your backend executes them, returns exact fixes to the agent when something is missing, and records a receipt for every decision. Learn more at usesalus.ai, explore the product, or read how Salus handles security and data.