IT security intelligence. Since 2006.Cloudflare services
TRUST-IT / Practical guide

Enterprise AI-agent security assessment before production

Decide whether an AI agent is ready for production using evidence about permissions, data access, tool execution, human approvals and recovery. A successful demonstration is only the start of the assessment.

Engineers reviewing an AI workflow and approval process
About 6 min read

Define the agent’s operational boundary

List the business task, users, data sources, tools, permitted effects and actions that must remain human decisions. Distinguish an assistant that drafts a response from an agent that sends it, changes a record or initiates a workflow. Document the identity used at each boundary: user, agent service account, connector and downstream application.

Create a versioned assessment baseline covering model configuration, instructions, retrieval policy, tool schemas and integration permissions. Name the system owner and the authority for accepting residual risk. The security decision should attach to a defined release and environment; changing a connector or granting a new tool can invalidate earlier results even when the visible chat interface looks unchanged.

Enforce data access at retrieval and execution

A knowledge assistant should retrieve only material the requesting user is authorised to use. Evaluate document permissions, group membership, tenant context and cached results when information is retrieved and delivered. Test whether a revoked entitlement stops subsequent access. A prompt that asks the model to respect permissions cannot replace enforcement by the retrieval and application layers.

Separate the agent’s technical identity from the user’s business authority. Minimise connector scopes and avoid a shared high-privilege credential for every task. Test cross-user and cross-tenant cases using synthetic records, including a valid user requesting a resource outside their scope. A convincing answer with an unauthorised source is still a security failure.

Treat external content as data with uncertain trust

An email, retrieved document, website or tool result may contain instructions that conflict with the authorised task. Prompt-injection testing should cover these indirect inputs as well as direct user prompts. Evaluate the actions and disclosures that follow, not only whether the model produces a warning or refuses an obvious test sentence.

A practical assessment uses controlled scenarios with harmless markers and synthetic data. Check whether a document can cause an unexpected tool call, alter a destination or introduce an unsupported instruction into a later workflow. Content filters are one layer; independent limits on permissions, outbound destinations and allowed effects reduce the consequences when model behaviour is wrong.

The control path for a consequential action

Illustrative control model: the language model proposes an action; application controls decide whether that exact action may execute. Apply the same policy to retries, background jobs and direct tool calls.

  1. Identify

    Resolve the user, tenant and delegated authority.

  2. Retrieve

    Return only sources the current identity may access.

  3. Propose

    Create a typed action with explicit parameters.

  4. Authorise

    Check policy and, where required, bind human approval.

  5. Execute & record

    Enforce the approved scope, log the result and prevent duplicate effects.

Bind approvals to a specific proposed action

For a material external action, the approver should see the exact destination, data, purpose and consequence. Approval needs a defined lifetime and must be tied to the action actually executed. If parameters change after approval, require a new decision. Avoid treating a broad chat acknowledgement as permission for a sequence of undisclosed later actions.

Tool handlers should validate typed inputs, authorised resources and permitted transitions independently of generated text. Consider how repeated requests, retries or a delayed callback could duplicate an effect. Define cancellation and recovery behaviour, and test whether a denied or expired approval reliably prevents execution rather than merely hiding a button in the interface.

Build an evaluation set with acceptance criteria

Combine ordinary user tasks, misuse cases and operational failures. Include a legitimate action that must succeed, an out-of-scope request that must fail, changed permissions, unavailable tools, malformed tool responses and conflicting retrieved material. Repeat selected scenarios where nondeterministic behaviour could change the result, while recording configuration and evidence for each run.

Agree acceptance thresholds before reviewing results, and separate safety-critical failures from quality defects. Report the tested population and untested combinations; passing a small set does not measure a universal safety percentage. Preserve test inputs, relevant outputs, permission decisions, approvals and tool results, with sensitive values minimised. Re-run affected tests after changes to policies, tools, prompts, models or retrieval.

Make observation and recovery part of readiness

Logs should connect a user request to retrieval decisions, proposed actions, approval and downstream outcome using stable correlation identifiers. Restrict access and avoid storing secrets or unnecessary document contents. A useful incident record explains what happened without creating a second repository of sensitive information.

Exercise a stop procedure for new tasks, a method to revoke connector access and a plan for partially completed workflows. A software rollback cannot automatically undo an email already sent or a record changed in another system. Identify compensating actions and responsible owners, then verify recovery with an agreed scenario. Monitor the permissions and integration health on which the release decision depends.

Turn an impressive demonstration into release evidence

On a small screen, scroll the table sideways to compare all columns.

Turn an impressive demonstration into release evidence
Test caseExpected controlEvidence for the release owner
A retrieved document asks for an unauthorised tool action.Source text cannot grant privileges or override execution policy.Test input, tool decision and outcome tied to the same trace.
A user requests another tenant’s information.Retrieval, cache and tool execution enforce tenant and record permissions.Denied attempts and legitimate same-tenant requests both tested.
The action changes after approval.Approval for old parameters cannot authorise the new action.Bound action digest or equivalent reference, approver and expiry.
The tool times out after committing a change.Retry handling reconciles state before repeating side effects.Recorded operation identifier and duplicate-prevention result.
A connector or model version changes.Relevant regression cases run against the new release baseline.Versioned evaluation results, open risks and rollback decision.

Turn the assessment into a release decision

A decision pack should contain the boundary diagram, identity and tool inventory, test catalogue, evidence, unresolved findings, operational playbooks and approval conditions. Distinguish “ready within the tested scope”, “ready with explicit restrictions” and “not ready”. The business owner should understand the restrictions that must survive deployment.

TRUST-IT can assess an existing agent, review a planned architecture or define security acceptance criteria for an AI pilot. The engagement connects technical evidence with business authority and operational responsibility. Our controlled-AI-agent technical example provides a deeper illustrative architecture; the assessment checklist helps an enterprise decide what evidence to request for its own implementation.

Further reading

Put the guidance to work

Assess prompt injection, data exposure, agent permissions, model behaviour, and the governance of AI adoption.

AI security & governance

What’s your next
technology challenge?

TRUST-IT / FIND YOUR NEXT STEP

How can we help?

Popular topics

Search the public TRUST-IT website.