Enterprise AI buyers have learned to ask better questions than they asked two years ago.
They ask whether the model is accurate. They ask whether it is aligned. They ask whether private data is protected, whether malicious prompts are blocked, whether outputs are filtered, whether logs exist, and whether the vendor can show a responsible AI policy.
Those questions matter. They are also incomplete.
Once an AI system can retrieve private data, invoke tools, alter records, send communications, issue refunds, initiate financial activity, modify production systems, operate machinery, or direct other agents, the decisive question is no longer only whether the system gives a safe answer.
The decisive question is this:
Who gave this system authority, what exact authority was granted, and what technically prevents it from exceeding that authority?
Capability is a technical fact. Authority is a human and institutional grant. The ability to identify, recommend, plan, or technically perform an action does not establish permission to execute it.
Enterprise AI safety begins to become real when buyers stop treating those two ideas as the same.
Capability Is Not Authority
A system can know what should be done without being permitted to do it.
That distinction is familiar outside artificial intelligence. A junior analyst may know that a customer record is wrong but lack authority to alter the system of record. A physician may recommend a treatment, but a pharmacy, insurer, hospital policy, and patient consent still shape what happens next. A security researcher may identify a vulnerability without receiving permission to exploit a production environment.
AI does not erase that structure. It makes the structure easier to accidentally bypass.
An enterprise assistant may be capable of reading a contract, finding a billing discrepancy, drafting a refund, composing a customer email, updating a CRM record, and opening a finance workflow. Those are capabilities. They become authority only when the organization has made a separate decision that this system may take those actions under defined conditions, for defined users, with defined evidence, and within enforceable technical limits.
Vendors often demonstrate capability because capability is impressive. A buyer should ask where the authority comes from.
Who approved autonomous execution? Which role approved it? Which policy records the delegation? Which tool calls are available? Which data can be reached? Which actions require confirmation? Which actions require an independent human? Which actions are impossible because the architecture withholds the permission entirely?
If the answer is mainly that the model has been instructed not to do the wrong thing, the answer is not enough.
Recommendation Is Not Execution
Recommendation and execution should be separate stages, especially when the result is consequential.
An AI system may recommend that a refund be issued. That does not mean it should be able to issue the refund. It may recommend that a production service be rolled back. That does not mean it should be able to deploy the rollback. It may recommend that a customer receive a notice, that an invoice be voided, or that an employee's access be changed. None of those recommendations should become execution merely because the system can call a tool.
The enterprise control question is not whether the model understands the task. It is whether a separate authorization layer evaluates the action before it happens.
That layer should know the requesting user, the user's role, the policy governing the action, the data involved, the financial or operational impact, the required approvals, the current state of dependencies, and the exact operation being requested. It should also be able to say no in a way the model cannot override.
This is why behavioral instructions and architectural controls are different things.
A prompt that says "do not send external email without permission" is useful only as a behavioral instruction. A mail integration that cannot send until a separately recorded approval token exists is an architectural control. A prompt that says "do not edit production records" is a request. A tool identity that has no write permission is a boundary.
Security-critical boundaries should exist outside the model whenever possible. The model may help prepare an action, but it should not be the final authority on whether the action is allowed.
Credentials Are Not Authorization
Possessing credentials is not the same as being authorized to use them for every available action.
This distinction becomes fragile in agentic systems. A platform may connect to enterprise tools through a service account. That service account may be powerful because it was convenient to integrate. The agent may then appear to act "for the user" while actually using a privileged platform identity with access the user does not have.
That is a quiet authority transfer.
An agent acting for a user should not silently inherit a higher-privileged platform identity. Identity and authority should remain preserved from the requesting person through retrieval, agent planning, tool selection, approval gates, execution, logging, and final outcome.
If a user is not allowed to see a document, the agent should not retrieve it. If a user is not allowed to change a record, the agent should not change it. If a downstream tool receives a request from an agent, the tool should know whose authority is being exercised and which scope was granted for that exact action.
OWASP's guidance on Excessive Agency makes the same point in practical terms: agent extensions should be limited to necessary functions, permissions should be minimized, high-impact actions should require approval, and authorization should be enforced in downstream systems rather than left to the language model's judgment.
That is not anti-AI. It is ordinary access control applied to systems that can now talk, plan, and act.
Logs Are Not the Same as Governed Transactions
Many vendors say they log conversations. That may be useful. It is not the same as being able to reconstruct a governed transaction.
If an AI system takes a consequential action, an auditor should be able to reconstruct the chain from request to outcome.
Who requested the action? Which identity was used? What data did the system retrieve? Which policies were evaluated? Which version of the model and tool policy was active? What did the system propose? What authorization was required? Who approved it? Was approval independent? What exact tool call executed? What changed? What notification was sent? What failed? What was withheld? What evidence proves the result?
A transcript alone rarely answers all of that.
Governance needs transaction evidence, not just conversation history. The evidence should show authorization decisions, not merely model words. It should survive after the session ends. It should be difficult for the acting system to rewrite. It should integrate with the organization's ordinary incident response, records, compliance, and security monitoring functions.
The F5-hosted NSS Labs paper usefully emphasizes observability, forensics, delegated authority, identity-context preservation, degraded operation, and independent validation as buyer concerns. That framework is valuable even though a paper hosted by a vendor should not be read as proof that any vendor satisfies its own criteria. The buyer still has to test.
Degraded Operation Reveals the Real Policy
Normal operation is the easiest condition under which to look safe.
The more revealing question is what happens when dependencies fail.
What happens if the identity provider times out? What happens if the retrieval index is stale? What happens if the policy engine cannot be reached? What happens if the audit pipeline is unavailable? What happens if the approval service is down? What happens if one agent in a chain cannot verify the authority state passed by another agent?
Under degraded conditions, a system's true policy becomes visible.
If consequential authority silently continues when identity, policy, retrieval, or audit dependencies fail, the system has chosen availability over governance. That may be acceptable for low-risk assistance. It should not be acceptable for actions that alter records, move money, contact customers, change permissions, deploy software, affect safety, or create legal obligations.
Fail-closed does not always mean the system must stop everything. A mature system can degrade by reducing authority. It may continue answering general questions while withholding private retrieval. It may draft a proposed action while refusing execution. It may allow read-only access while disabling writes. It may preserve a work item for later approval rather than acting without evidence.
The point is not theatrical shutdown. The point is predictable loss of authority when the conditions for authority cannot be proven.
Agent Chains Must Not Propagate Authority by Accident
Multi-agent systems create a second authority problem.
Information may need to move from one agent to another. Authority should not automatically move with it.
An intake agent may summarize a customer issue. A diagnostic agent may inspect permitted logs. A billing agent may calculate a possible refund. A communications agent may draft a response. Each handoff should preserve the user's identity, the task purpose, the data classification, the approved scope, the allowed actions, and the remaining approvals required.
The receiving agent should not treat the previous agent's output as a new grant of authority. It should not assume that because one agent could retrieve information, another agent may disclose it. It should not assume that because one step was permitted, the combined workflow is permitted.
Several individually allowed actions can combine into an unauthorized result. A system may be allowed to read a support ticket, draft a message, and update a customer record. Combined in the wrong order or under the wrong identity, those actions can disclose private information, waive a contractual right, or create a false business record.
This is why chain analysis matters. Buyers should ask whether the vendor can detect unsafe combinations of otherwise permitted actions and whether authority is scoped to the whole workflow, not just to individual tool calls.
Ten Questions Enterprise Buyers Should Ask
Before buying or expanding an agentic AI system, ask these questions plainly and require evidence.
1. What actions can the system technically perform, including through tools, plugins, APIs, agents, connectors, and inherited platform capabilities?
2. Which of those actions may it perform autonomously, under what conditions, and who approved that delegation?
3. Are forbidden actions architecturally unavailable, or are they merely discouraged by prompts, policies, or model instructions?
4. When the agent acts, does it inherit the requesting user's permissions, a constrained delegated scope, or a privileged system identity?
5. Can several permitted actions combine into an unauthorized outcome, and does the system analyze action chains before execution?
6. Which actions require explicit confirmation, independent approval, or multiple-human authorization before execution?
7. What happens to authority when identity, policy, retrieval, tool, approval, logging, or audit dependencies fail?
8. Can authority be revoked during an active workflow, and what happens to queued, partially completed, or downstream actions after revocation?
9. Can each consequential action be reconstructed from user request through retrieved context, policy evaluation, authorization, execution, and final outcome?
10. Does the vendor permit realistic independent testing under adversarial conditions, long-running workflows, multi-agent chains, and degraded dependencies?
A vendor that can answer these questions with architecture, logs, tests, and contractual commitments is in a different posture from a vendor that answers with trust language.
Buyers should notice the difference.
Refusing Authority Is a Safety Mechanism
AI safety is often framed as the model's ability to refuse a harmful prompt. That matters. A system that declines dangerous instructions is better than one that follows them.
But enterprise safety cannot stop there.
The most important safety mechanism may not be the model's ability to refuse a harmful prompt. It may be the organization's ability to refuse the model authority.
Refuse authority to reach data the user cannot reach. Refuse authority to use credentials beyond the user's scope. Refuse authority to execute where recommendation is enough. Refuse authority to continue when governance dependencies fail. Refuse authority to let one agent's output become another agent's permission. Refuse authority to hide behind a generic platform identity. Refuse authority to act when the evidence cannot be reconstructed.
This is not a rejection of intelligent systems. It is how intelligent systems become usable inside responsible institutions.
The enterprise buyer's job is not to be dazzled by what the system can do. The buyer's job is to determine what the system is allowed to do, who made that decision, and what prevents the boundary from moving on its own.
Who gave the system authority, and can you prove that it has no more authority than that?
