What Changes When AI Is in the Product
A standard web application pentest evaluates authentication, input validation, business logic, and server-side code. Testing an AI feature adds a separate layer: how the model responds to adversarial prompts, whether a retrieval pipeline can leak documents it shouldn't, whether a connected tool or API can be tricked into an unauthorized action, and whether the AI inherits more trust than it should from the systems around it. Firms differ in how much of that layer they actually cover and how they scope it, which is the main thing worth comparing before picking one.
With that gap in mind, the question becomes which firm actually covers it, and how. Some have built out a named AI/LLM testing methodology with its own framework mapping. Others fold AI-specific work into a broader application test or a wider security practice. The table below is a starting point for comparing how each firm scopes that work, followed by a closer look at each.
Quick Comparison
| Provider | Who it fits | AI-specific coverage | Where it stands out |
|---|---|---|---|
| Compass IT Compliance | Teams that also need compliance work (SOC 2, CMMC, PCI, ISO 27001) alongside testing | AI and LLM penetration testing, plus AI governance work | One partner for testing and compliance, national reach, decade-plus track record |
| Software Secured | SaaS and AI-native teams wanting a dedicated AI/LLM pentest product | Purpose-built AI Pentesting: models, RAG, MCP tools, agents, AIwritten code | Explicit AI framework mapping (MITRE ATLAS, OWASP LLM Top 10) and built-in retesting |
| Redbot Security | Teams wanting a small, senior-led boutique over a larger firm | AI & LLM Security Testing covering prompt injection, RAG, agents, connected trust paths | Boutique model with ISO 27001:2022 and SOC 2 Type I/II assurance |
| Packetlabs | Teams that want CREST accreditation and a heavily manual testing model | AI/LLM Penetration Testing: injection, leakage, plugin risk, business logic abuse | 100% OSCP-minimum staff and CREST accreditation, per the firm |
| Raxis | Teams wanting a U.S.-based firm with public CVE research behind its testers | Eight-step AI/LLM methodology mapped to OWASP LLM Top 10 and MITRE ATLAS | Published CVE discoveries and detailed public methodology |
| Virtue Security | Teams wanting a small NYC boutique focused only on app and API testing | Pentests tailored to AIenabled applications | 15+ years average team experience, narrow application-only focus |
| TrustedSec | Teams wanting AI security bundled with governance and incident response | AI Security Assessment, offensive AI Red Team, governance advisory, AIbreach IR | CREST certified, broad consulting bench beyond pentesting alone |
| COE Security | Teams wanting AI testing bundled with many adjacent services under one vendor | Defined AI/LLM methodology mapped to NIST AI RMF and OWASP LLM Top 10 | Wide service catalog spanning IoT, firmware, blockchain, and AI |
A Closer Look at Each Firm
1. Compass IT Compliance
Compass is a national IT security and compliance consulting firm founded in 2010 and based in Rhode Island, serving more than 1,000 clients across financial services, healthcare, higher education, retail, technology, and other sectors.
Compass IT Compliance performs AI and LLM penetration testing, evaluating how a product's prompts, model behavior, and connected integrations hold up against adversarial input. Its risk and business resiliency practice separately covers AI governance work for organizations building internal AI policy. Beyond AI, its penetration testing practice covers web application, network, wireless, mobile, cloud, and social engineering testing, built on OWASP, OSSTMM, and NIST methodologies with black, gray, and white box options and a formal rules-of-engagement process. Testing is positioned as hands-on and expert-led rather than scan-andreport, with immediate notification for high-risk findings and reporting split between a technical walkthrough and an executive summary.
Because Compass also runs SOC 2, PCI DSS, HIPAA, CMMC, and ISO 27001 practices under one roof, teams building AI products that also need to satisfy a compliance framework can fold that work into the same engagement instead of coordinating between separate testing and audit vendors. Compass is a long-term partner rather than a one-time vendor, with a bench of specialists rather than a single consultant handling an account.
2. Software Secured
Software Secured is a manual penetration testing firm founded in 2010 by Sherif Koussa and based in Ottawa, Canada, built around bringing bank-grade application security to growing software companies. The firm reports more than 2,000 pentests over the last five years and more than 350 high-growth SaaS startups, scaleups, and SMBs as clients.
Its AI Pentesting service targets LLMs, agents, and MCP servers specifically, covering model behavior, RAG data retrieval, connected tool and MCP integrations, agent workflows, and AI-written code, with test plans mapped to the MITRE ATLAS Matrix, Google's SAIF risk framework, and the OWASP Top 10 for ML. Findings are also mapped to OWASP LLM Top 10, ISO 42001, GDPR Article 32, and SOC 2 criteria to support compliance reviews. The firm reports 0 false positives, includes built-in retesting, and prices its AI pentesting starting at $10,800.
3. Redbot Security
Redbot Security is a boutique offensive security firm that launched in 2018, describing itself as intentionally built around senior-led delivery over high assessment volume. The firm holds ISO 27001:2022 certification and maintains SOC 2 Type I and Type II assurance, with operators holding certifications including OSCP, CRTO, GPEN, and CISSP.
Its AI and LLM Security Testing service covers prompt injection and jailbreak attempts, RAG and retrieval poisoning, agent and tool abuse, and the identity and cloud trust paths that connect an AI system to the rest of an environment. The firm frames its approach as testing the full AI system rather than just the model, with every finding manually validated through proof-of-concept evidence.
4. Packetlabs
Packetlabs is a penetration testing firm headquartered in Toronto with outposts in San Francisco, Calgary, and Sydney, describing itself as CREST-accredited and SOC 2 Type II attested. The firm reports that all of its staff hold at least an OSCP certification and that 95% of its testing is manual rather than automated scanning.
Its AI/LLM Penetration Testing service covers prompt injection and jailbreak attempts, data leakage and cross-tenant exposure, model and API authentication and rate limiting, identity and access abuse through AI workflows, insecure plugin and integration risk, and business logic abuse where AI flaws chain into application or infrastructure weaknesses. The firm publishes a direct comparison of AI/LLM testing against standard application testing on its site.
5. Raxis
Raxis is a penetration testing firm based in Atlanta that positions itself around human-led, U.S.-based testing, citing published CVE discoveries in enterprise software by its own engineers. The firm also offers penetration testing as a service for continuous coverage between annual engagements, alongside physical and OT testing.
Its AI & LLM Penetration Testing service follows an eight-step methodology covering architecture review and threat modeling, system prompt and configuration analysis, adversarial prompt testing, data extraction testing, RAG and retrieval pipeline assessment, agent and tool exploitation, output validation, and reporting. Findings are mapped to the OWASP Top 10 for LLM Applications and MITRE ATLAS.
6. Virtue Security
Virtue Security is a boutique penetration testing firm based in New York City, founded by Elliott Frantz, focused solely on application, API, and cloud penetration testing rather than a broader security consulting practice. The firm reports more than 15 years of average team experience and describes its assessments as manual and tailored to each client's technology, with dedicated retesting included.
Virtue Security lists AI as one of its focus industries alongside fintech and healthcare, tailoring pentests specifically for AI-enabled applications to help organizations stay ahead of emerging risk in that area.
7. TrustedSec
TrustedSec is a cybersecurity consulting firm founded by David Kennedy in 2012 and based in Fairlawn, Ohio, offering penetration testing alongside incident response, security program design, and advisory services. The firm holds CREST certification and describes itself as practitioner-led and adversary-informed.
Its AI Cybersecurity Consulting practice includes an AI Security Assessment of systems, pipelines, and integrations; offensive AI Red Team testing covering prompt injection, model manipulation, and abuse-case simulation; AI governance and advisory work; and incident response specifically for AI-related breaches. The firm's leadership describes its approach as using AI to support consultants rather than replace the judgment behind an engagement.
8. COE Security
COE Security is a cybersecurity services firm that traces its origins to 2007 and describes a decade-plus history expanding from application testing into cloud, IoT, blockchain, and AI security services, including a proprietary AI vulnerability identification tool launched in 2024.
Its AI & LLM Penetration Testing service follows a defined process covering scope and component definition, attack surface enumeration, prompt injection and manipulation testing, output filtering and alignment validation, training data exposure risk assessment, plugin and API abuse testing, authentication and session review, and adversarial input analysis. The firm states its methodology aligns with the NIST AI Risk Management Framework and the OWASP LLM Top 10, and includes post-remediation retesting.
Red Flags Worth Watching For
- A vendor that cannot describe what a human tester actually does versus what a scanner does.
- No sample report available before you sign.
- Retesting billed as a separate, optional line item rather than included in scope.
- AI testing pitched as a checkbox add-on with no framework (OWASP LLM Top 10, MITRE ATLAS) behind it.
Questions Worth Asking Before You Sign
A few questions tend to separate a firm that has actually built out AI testing from one that added it to a services page recently.
- Scope: what gets tested beyond the model itself, including retrieval sources, connected tools and APIs, and the identity boundaries around them.
- Delivery: whether testing is performed by staff or contractors, and how much of the process is manual versus automated.
- Reporting: for a sample report to see whether findings come with reproduction steps and business impact context, not just a severity score.
- Retesting: whether retesting after remediation is included or billed separately.
- Compliance: whether the firm can map findings directly to a compliance framework if a deadline is driving the engagement.
Frequently Asked Questions
Is testing an AI product different from a regular application pentest?
Yes, in scope. It typically includes everything a standard application pentest covers, plus testing specific to how the AI system handles prompts, retrieval, connected tools, and agent behavior.
Does an AI pentest replace ongoing monitoring?
No. A pentest is a point-in-time assessment. Teams shipping AI features regularly should expect to retest after significant changes to the model, prompts, or connected tools, not just once a year.
Is it fine to rely on automated AI security scanning to save money?
Automated scanning has a role, but for AI-specific risks like prompt injection and agent misuse, purely automated tools tend to miss business-logic issues that require a human tester reasoning about how the system could be manipulated.
What does this typically cost?
Pricing varies by scope and firm, but published starting points for dedicated AI pentesting services run in the five-figure range for a focused engagement, with broader scopes priced higher.