Pricing

Straightforward Pricing for AI Reliability

One-time assessments for pre-launch teams. Quarterly retests when things change. Project-based builds for teams that need infrastructure or compliance work done properly. No ongoing contracts.

Assess

$499/

one-time

A complete point-in-time test of your AI system. Red team attack run plus quality baseline, packaged as a severity-ranked report in 48 hours. No contract.

  • Red Team Assessment — 374 attack vectors across AI and web layers
  • Quality Baseline — 64 eval cases across 6 quality dimensions
  • Severity-ranked HTML + PDF report with full payload and response evidence
  • 48-hour turnaround from kick-off call
  • One retest cycle included after your team remediates findings
Book a Free Scoping CallDownload a real sample report
Most Popular

Retainer

$299/

per quarter

$499/qtr adds Eval Baseline re-run + 30-min findings review call

Scheduled re-runs after model updates, prompt changes, or remediation cycles. Keeps your findings current without a continuous monitoring contract.

  • $299/qtr — Re-run all 374 red team vectors after any change. Delta report shows exactly what shifted vs. your last run.
  • $499/qtr — Everything above plus a full Eval Baseline re-run (64 cases, 6 dimensions) and a 30-min findings review call.
  • 48-hour priority turnaround
  • Cumulative findings history across every cycle so you can track improvement over time
  • Book on a fixed quarterly schedule or ad hoc after any model swap, prompt update, or remediation
Book a Free Scoping Call

Enterprise

Custom

scoped on a call

Project-based engagements for teams that need infrastructure built or compliance documented — not just tested.

  • Eval Pipeline Build — from $3,000 one-time: LangFuse or LangSmith configured on your stack, eval suite built, runbook and handover session included. You own everything after.
  • Full project price: $3k–$8k depending on scope. Standard LangFuse/LangSmith setup at the low end — custom eval suites, RAG pipeline integration, or multiple environments at the high end.
  • Compliance Gap Analysis — $999–$2,499 one-time: structured review against HIPAA, SOC 2, or EU AI Act. Gap report with prioritised remediation items and audit-ready documentation.
  • Custom programs for multi-system AI teams and formal governance requirements
Book a Free Scoping Call

Every Engagement

What You Can Always Expect

📋

Evidence-backed reports

Every finding includes the exact payload sent, the response received, and a severity rating. Not a summary — a full audit trail.

🔁

Retest after remediation

We retest every finding after your team has addressed it to confirm resolution. Verification, not just discovery.

🤝

NDA on request

We sign NDAs before any engagement starts. Your system architecture and findings stay confidential.

48-hour first turnaround

From kick-off call to first findings report. No long pre-engagement discovery cycles before you see results.

Questions

Common Pricing Questions

What happens on the free scoping call?

The free 30-minute scoping call is where we review your AI stack, identify your highest-risk areas, and walk you through what an engagement would look like. It is not a full assessment — it is a no-commitment way to understand the scope before you decide.

Can I start with Assess and add a retainer later?

Yes, and it's a common path. Your initial assessment baseline becomes the reference point for every future retest — each delta report shows exactly what changed since your first run. You can add the quarterly retainer at any time after your first assessment.

What access do you need to run an engagement?

For red teaming: API access to your AI endpoints (we use a secure cURL paste-and-parse workflow — no persistent access to your infrastructure). For eval pipeline builds: access to your CI/CD pipeline and deployment tooling, scoped specifically on the engagement call. We sign NDAs before anything is shared.

Do you work with startups or only enterprises?

We work with any team where AI reliability matters — from well-funded startups shipping their first LLM product to enterprise teams managing complex multi-agent systems. The Assess tier is specifically designed for teams who want rigorous testing without a long-term commitment.

Not sure which tier fits?

Book a free 30-minute scoping call. We'll understand your stack and tell you exactly what we'd recommend — no commitment required.

Book a Free Scoping Call →