AI Agent Production Readiness.
Your AI agents are in production. Are they production-grade?
We review one production or near-production agent deployment across identity, permissions, auditability, observability, cost controls, security exposure, and deployment posture—then deliver a written, prioritized fix plan within 2–3 business days after kickoff and complete access.
Start by email. We will confirm fit, scope, price, and the earliest available start date.
Audit flow
01
Evidence
Configuration, code, logs, and architecture context
02
Controls
Seven production-readiness domains
03
Findings
Gaps, consequences, and effort estimates
04
Priorities
P0, P1, and P2 remediation roadmap
Fixed fee
$1,500–$3,000
Written report
2–3 business days after kickoff and complete access
Low-friction review
Read-only access or screen-share
The production gap
A working agent is not necessarily a production-grade agent.
A successful demo proves that a path can work. Production requires evidence that the system remains controlled when permissions expand, tools fail, costs rise, inputs turn hostile, or a rollback becomes necessary.
01 · Identity
Authentication and identity
OAuth 2.1, enterprise identity and single sign-on, token handling, and secrets management.
02 · Permissions
Authorization
Tool-level permissions, least privilege, and controls that limit scope creep and unauthorized actions.
03 · Trace
Auditability and traceability
Tool-call logging, action reconstruction, and retention sufficient to understand what the agent did.
04 · Signal
Observability
Errors, latency, alerts, and agent or tool failures that would otherwise remain silent.
05 · Guardrails
Cost and abuse controls
Rate limits, spend caps, and runaway-loop protection for predictable operating cost.
06 · Exposure
Security exposure
Prompt injection, tool poisoning, data-exfiltration paths, and tenant-isolation risks.
07 · Recovery
Deployment posture
Hosting controls, protocol currency, release safety, and rollback capability when production changes fail.
The practical consequences: unauthorized actions, missing traceability, silent failures, uncontrolled cost, data exposure, and fragile rollback.
Concrete output
A fix plan your technical team can act on.
The deliverable is not a generic maturity score. It is a written record of the architecture as found, the production gaps, and the order in which to address them.
Primary deliverable
8–12 page written report
Findings grouped by audit domain, with enough context to understand the condition, operational consequence, and recommended response.
Executive risk snapshot
One non-technical page for leadership.
Architecture as found
A shared view of the reviewed deployment.
P0 / P1 / P2 roadmap
Fix now, address within 30 days, or place in backlog.
Effort estimates
A practical size for every remediation item.
Quick wins
Improvements expected to require less than one day each.
30-minute readout
A separately scheduled walkthrough of findings and next steps.
Recommended next engagement
One appropriately sized and priced option when the evidence supports follow-on work.
How it works
A focused review, from context to priorities.
01
Start by email
Send the deployment context to PLEESH. We respond to a new audit inquiry within one business day.
02
Confirm fit and scope
We qualify the deployment by email and may arrange a 20-minute fit-and-scope call. That conversation is not the paid technical audit.
03
Agree the start date
After written scope confirmation, cleared upfront payment, and confirmation that access will be ready, we agree a start date with at least three business days’ lead time in Central Time (Dallas).
04
Kick off and review
Hold a 60–90 minute kickoff with the deployment owner, then review read-only evidence or work through it by screen-share across the seven domains.
05
Receive the roadmap
The written report is delivered within 2–3 business days after kickoff and complete, usable access. The 30-minute readout is scheduled separately.
Scope and pricing
One audit. A scope-dependent fixed fee.
$1,500–$3,000
Scope is confirmed before invoice. Payment is due in full before kickoff. PLEESH schedules no more than one active audit at a time.
Scope examples—not packages
$1,500 endpoint
One environment and no more than three agent integrations.
$3,000 endpoint
Up to two environments and no more than eight integrations.
The confirmed fee depends on the environment count, integration count, evidence, and access approach. It does not cover unlimited systems.
Clear boundaries
What this audit is—and is not.
It is
An evidence-based review of one agent deployment and a prioritized remediation plan grounded in the seven audit domains.
The work identifies production gaps, explains their operational consequences, estimates remediation effort, and separates urgent fixes from backlog work.
It is not
- Implementation of the findings
- A penetration test, certification, or compliance audit
- Model or prompt-quality evaluation
- New feature development
- Broad AWS, CI/CD, or general infrastructure implementation
After the audit
Diagnose first. Implement only what the evidence supports.
The audit is the front door. Any implementation is optional, separately scoped, and sized from the findings.
Path 01
Agent and Model Context Protocol (MCP) hardening
Repair identity, authorization, tool behavior, traceability, and agent-layer guardrails.
Path 02
Infrastructure and DevOps remediation
Address the hosting, delivery, observability, release, and rollback controls beneath the agent.
Path 03
Ongoing agent operations
Maintain operating discipline as integrations, permissions, costs, and production behavior change.
About PLEESH
AI infrastructure and reliability, applied to the agent layer.
PLEESH is a global AI infrastructure and reliability consultancy with a presence in Dallas, Texas.
We help teams turn production and near-production AI agents into systems that are secure, observable, auditable, cost-controlled, and operationally resilient.
Our approach applies site reliability engineering (SRE) and DevOps operating discipline to the agent, tool, and cloud layers. We begin with a focused production-readiness audit, then help implement the highest-priority fixes when needed.
Capability areas
- AI agent and Model Context Protocol (MCP) production readiness
- OAuth 2.1, enterprise identity/SSO, tokens, and secrets
- Tool authorization, least privilege, validation, and safe behavior
- Audit logging, observability, alerting, and incident readiness
- Rate limiting, spend controls, and runaway-loop protection
- Prompt-injection, tool-poisoning, exfiltration, and tenant-isolation exposure reviews
- AWS, Kubernetes, Terraform, Docker, continuous integration and delivery (CI/CD), deployment, and rollback design
- Python, TypeScript, Node.js, PostgreSQL, APIs, and custom integrations
FAQ
Questions before you request an audit.
The audit is designed for a real deployment with a technical owner—not an idea that has not yet reached implementation.
Who is the audit for?
Teams with one production or near-production agent deployment and a technical owner who can join the kickoff. Typical environments use Claude, ChatGPT, Copilot, a custom agent, MCP, or a comparable tool layer.
What access is required?
Read-only access where possible, or a screen-share walkthrough. Useful evidence includes agent and tool configuration, relevant code or infrastructure as code, representative logs, and an architecture sketch if one exists. Do not email credentials, secrets, production logs, or confidential evidence with the initial inquiry.
What if our architecture is not documented?
Formal documentation is not required. Existing notes and a walkthrough are enough to establish the architecture as found.
Will this disrupt production?
The audit is a review, not an implementation engagement. PLEESH prefers read-only evidence or screen-share and does not make production changes as part of the audit.
Is this a penetration test or SOC 2 audit?
No. It is an evidence-based production-readiness review and remediation plan. It is not penetration testing, certification, compliance work, or a guarantee that a deployment is safe.
Do you review prompts or model quality?
No. Model and prompt-quality evaluation are outside this audit. The review focuses on the production controls around the agent and its tools.
Can you implement the findings?
Potentially. Implementation is optional, recommended only when supported by the findings, and scoped separately after the audit.
When does the 2–3-business-day clock start?
After the kickoff and after complete, usable access or evidence has been provided. The 30-minute readout is scheduled separately and is not promised within that delivery window.
Start with evidence
Find the production gaps before they become production incidents.
Email PLEESH to confirm fit, scope, fixed price, and availability. We respond to new audit inquiries within one business day.