AI application security testing

What can your AI feature read, reveal or do?

Testing for AI applications, retrieval systems and agents, focused on data boundaries.

Human-led testing.
Manually verified findings.

  • OSCP-certified testers
  • Written testing boundaries
  • Actionable remediation
What we investigate

A scope with a purpose.

Prompt injection

Direct and indirect instructions through user input and retrieved documents.

Data disclosure

Cross-user retrieval and the authorisation checks around data sources.

Tools and downstream actions

Tool permissions, approval boundaries and unsafe handling of generated output.

Before testing

What we need from you

  • Application architecture, model integrations and retrieval sources.
  • The tool inventory and the role matrix.
  • Synthetic sensitive data and sandboxed actions.
  • Repeatable test cases for non-deterministic behaviour.
Use the scoping checklist
After testing

What you can act on

  • Application-level reproduction scenarios.
  • The prompts and tool traces involved.
  • Observed boundary failures with regression-test suggestions.

Reporting and retest terms are agreed in writing. The record separates verified fixes from outstanding work.

Limits matter

What this test does not cover

  • A guarantee against invented output or every jailbreak.
  • General model quality and training-data audits.
  • Destructive tool actions.
  • Attacks on the model provider.
Questions before you book

Practical answers.

Is this just trying prompts against a chatbot?

No. A prompt matters when it crosses an application boundary, reveals restricted data or causes an unauthorised action. We examine the identity, retrieval and tool controls around the model.

What does a penetration test cost?

From A$7,500 ex GST for one web application with its API and two user roles. That covers five testing days, the report and a retest of critical and high findings. More applications, endpoints or cloud accounts give an indicative range. The price is fixed once scope is agreed, in writing, before work starts.

How long will it take?

Testing effort and elapsed delivery time are different. We agree both after reviewing the scope, access readiness and your deadline. Leave time for remediation and a focused retest.

An AI feature still sits on ordinary infrastructure. The interface is web application penetration testing, the service behind it is API penetration testing, and the account holding the model and its data is cloud penetration testing. We will not claim a model has been made safe, and what a penetration test report should contain sets the standard we report against.

Let’s scope it

Let’s define your AI application test.

Share the assets and your reason for testing. We will confirm the approach, fee and schedule.

Request a quote

Last reviewed: