The Four-Percent Problem: Why Speed Kills Governance¶
Gartner projects that 40% of enterprise applications will embed task-specific AI agents by the end of 2026, a dramatic leap from less than 5% adoption in 2025. This surge is not just a trend; it is a structural shift in how software operates. As teams scramble to integrate these autonomous capabilities, they are frequently bypassing the traditional software procurement rigor that governed SaaS rollouts for the last decade. The result is a growing inventory of third-party AI agents that operate with permissions, data access, and financial exposure that few stakeholders fully understand.
When you buy a tool, you inherit its governance gaps. Unlike a static API, an AI agent makes decisions. It interprets intent, selects tools, and executes actions. If the vendor's architecture lacks robust identity management or spend controls, your organization absorbs the risk. This is not a hypothetical concern. The market is moving faster than the legal and security teams can keep up. You need a standardized way to evaluate these vendors before you sign the contract. This checklist provides the due-diligence questions you must ask to ensure the agent you deploy is governed, observable, and safe.
Identity and Permission Architecture¶
Agents are not users, but they must act as users. The first failure point in most agent deployments is identity management. If an agent uses a shared service account with broad permissions, you have created a single point of failure. You need to know how the vendor handles the digital identity of the agent itself.
- Does the platform assign stable, unique identities to each individual agent instance? Ensure that Agent A cannot inherit the permissions of Agent B through a shared token.
- Are permissions scoped and time-bound? The system should support expiring grants that automatically revoke access after a defined period, rather than relying on manual cleanup.
- How does the platform handle decommissioned identities? When an agent is retired or a project is closed, is the identity immediately invalidated, or does it linger in the system as a dormant risk?
- Can you enforce least-privilege access? The agent should only have access to the specific APIs and data stores required for its task, not broad read/write access to the entire environment.
If a vendor cannot demonstrate granular, per-agent identity control, you are effectively giving a robot your root password. This is a non-negotiable baseline for any enterprise deployment.
Human Approval Gates and Kill Switches¶
Autonomy is a spectrum, but your tolerance for error is likely at the conservative end. You must define which actions are allowed to proceed without human intervention. The most dangerous agents are those that can execute high-stakes actions — such as transferring funds, deleting records, or sending external communications — without a checkpoint.
- Which specific actions require explicit human approval? The vendor should allow you to define a list of "high-risk" actions that trigger a mandatory review step before execution.
- Are approval gates bypassable? You need to verify that no developer or administrator can configure the system to skip human review for critical operations. The gate should be hard-coded or enforced at the infrastructure level, not just at the application layer.
- Is there a global kill switch? If an agent begins looping or behaving erratically, you need a mechanism to stop all agent activity across the entire platform instantly. This should be a single action, not a multi-step process that requires accessing multiple dashboards.
- What is the latency of the kill switch? If the agent is in the middle of a transaction when you pull the switch, how quickly does the system halt? You need to know if the switch stops the next step or if it can interrupt the current execution cycle.
Ask your vendor to demonstrate the kill switch in a live environment. Do not take their word for it. If they cannot show you how to stop the bleeding in under ten seconds, they are not ready for production traffic.
Spend Controls and Financial Liability¶
AI agents consume resources. They make API calls, process tokens, and interact with paid services. Without strict financial guardrails, a logical error in the agent's reasoning can lead to catastrophic costs. There are reported cases of developers facing a single-incident API burn of over $700 due to infinite loops where the agent repeatedly called an external service without a termination condition. While $700 may seem manageable for a large enterprise, it represents a failure of control. Multiply that by the volume of agents you deploy, and the exposure becomes significant.
- Do you have per-agent budget limits? You should be able to set a hard ceiling on the total cost associated with a specific agent or project. If the agent exceeds this limit, it should stop operating.
- Are there hard ceilings on token usage? The platform should allow you to cap the number of tokens processed per hour or per day to prevent runaway loops.
- Who pays for a runaway loop? If the agent enters an infinite loop and consumes resources, does the vendor absorb the cost of the infrastructure, or is it passed through to you? Clarify this in the contract.
- Can you set alerts for anomalous spend? You need real-time notifications if an agent's consumption spikes unexpectedly. This allows you to intervene before the budget is exhausted.
Financial control is not just about saving money; it is about stability. A system that can halt itself when costs exceed a threshold is a safer system than one that relies on human monitoring to catch errors.
Audit Trails and Evidence Posture¶
In the event of an incident, you will need to reconstruct what the agent did, when it did it, and why. The vendor's logging capabilities determine whether you can perform a meaningful post-incident review. Vague or incomplete logs are a liability.
- Are logs immutable? Once an action is logged, it should be tamper-proof. If a vendor or administrator can edit or delete logs, the integrity of the audit trail is compromised.
- What is the export format? You should be able to export logs in a standard, machine-readable format such as JSON. This allows you to ingest the data into your own SIEM or data lake for long-term retention and analysis.
- What is the retention policy? How long does the vendor store logs? If your internal compliance requirements mandate a specific retention period, the vendor must support it, or you must have the ability to export and store the data yourself.
- Does the platform support EU AI Act evidence requirements? If you operate in the European Union, you must check your obligations regarding transparency and logging for high-risk AI systems. Ask the vendor specifically how their logging features align with these regulatory expectations.
Do not assume that because the vendor provides a dashboard, you have sufficient evidence. Dashboards are for human consumption; logs are for forensic analysis. Ensure you have access to the raw data.
Data Handling and Privacy¶
Agents need context to function. They may process customer data, internal documents, or proprietary code. You need to understand how this data is handled, stored, and potentially used for model improvement.
- Is your data used for training? Clarify whether the vendor uses your input data to train or fine-tune their models. If you do not want your data contributing to the public model, this must be explicitly excluded in the contract.
- Who are the sub-processors? The vendor may use third-party services for storage, inference, or monitoring. You need a current list of sub-processors and their locations to assess data residency risks.
- What happens to data at contract termination? You need a clear process for the deletion of all your data, including backups, once your contract ends. Ask for a certification of deletion.
- How are breaches notified? Define the timeline and method for breach notification. If a vulnerability is discovered in the agent's context window, how quickly will you be informed?
Data privacy is not just a legal issue; it is a trust issue. If your customers or employees know their data is being processed by an opaque third-party agent, you risk reputational damage. Transparency is key.
Commercial Terms and Exit Strategy¶
The vendor landscape is fragmented. You will see pricing structures ranging from a $29/month Core plan for individual developers to Enterprise tiers costing $2,000 to $2,499/month for advanced observability and support. This wide spread reflects different levels of service and capability. However, the cost is only part of the equation. You must consider the terms of engagement and your ability to leave.
- Is the pricing transparent? Look for hidden costs such as per-token overages, setup fees, or mandatory support contracts. The $29/month entry point is attractive, but ensure it covers the features you need without surprise add-ons.
- What are the exit and termination rights? If you decide to leave, how do you retrieve your configurations, workflows, and data? You should not be locked into a proprietary format that is difficult to migrate.
- Who owns the configurations? The prompts, workflows, and rules you build with the agent are intellectual property. Ensure you retain full ownership of these assets, even if you terminate the contract.
- Is there a notice period for price increases? You need protection against sudden, significant price hikes that could disrupt your budget. Look for clauses that allow you to negotiate or terminate if prices increase beyond a certain threshold.
Your exit strategy is as important as your entry strategy. A vendor that makes it difficult to leave is a vendor that has less incentive to maintain high service levels after you have signed.
The One-Week Assessment Plan¶
You do not need months to evaluate a vendor. You need a structured, one-week process that forces clarity. Here is a practical plan to execute this due diligence efficiently.
- Days 1-2: Send the Questionnaire. Distribute this checklist to the vendor's sales and technical teams. Ask them to answer in writing. Do not rely on verbal assurances. Written responses create a paper trail and force the vendor to be precise. If they cannot answer a question, they must state that clearly. Ambiguity is a red flag.
- Days 3-4: Review Evidence. Request demonstrations of the kill switch, spend controls, and audit logs. Ask for sample log exports. Verify that the permissions model works as described. If possible, run a small pilot project in a sandbox environment to test the agent's behavior under stress. Look for gaps between their marketing claims and their actual implementation.
- Day 5: Score the Responses. Create a simple scoring matrix. Assign a score to each section: Identity, Approval Gates, Spend Controls, Audit, Data, and Commercial Terms. A vendor that fails on Identity or Spend Controls should be disqualified immediately, regardless of how good their other features are. These are safety-critical features.
- Weekend: Decision Rule. Apply a strict pass/fail rule. Any vendor that cannot answer the kill-switch and spend-ceiling questions fails by default. If they cannot show you how to stop the agent or cap the costs, they are not ready for your production environment. Do not compromise on these two items.
This process is designed to be rigorous but fast. It shifts the burden of proof to the vendor. They must demonstrate that they have solved the governance problems, not just the AI problem.
Conclusion: The Governed Agent Mesh Playbook¶
Implementing this checklist is the first step, but you need a framework to manage the agents you deploy. The Governed Agent Mesh Playbook — Studio Edition ($29) provides the operational structure to keep you honest. It includes a vendor matrix that helps you track what each tool does and does not provide, ensuring you are not relying on a vendor that lacks critical safety features. Additionally, it offers approval and spend-limit policy templates that you can adopt immediately with your winning vendor. These templates translate your due diligence findings into enforceable internal policies. Do not leave governance to chance.
Use the playbook to build a mesh of agents you can actually govern. The Governed Agent Mesh Playbook — Studio Edition ($29) turns these twenty questions into enforceable policy — buying a vendor is a start; governing them is the job.