AI System Audit & Pressure Testing by Jeff PadgetAI System Audit & Pressure Testing by Jeff Padget
AI System Audit & Pressure TestingJeff Padget
Cover image for AI System Audit & Pressure Testing
Your AI system works in the demo.
Now let’s find out what happens when reality stops cooperating.
I audit and pressure-test AI agents, workflows, automations, memory systems, and multi-agent architectures to identify brittle assumptions, hidden failure modes, confusing boundaries, and places where a system can behave correctly according to one component while still failing as a whole.
The goal is not to prove that a system is bad.
The goal is to discover how it can fail while there is still time to do something useful about it.
Depending on the system, testing may examine:
• Agent roles and responsibility boundaries • Tool permissions and autonomy • Human approval checkpoints • Prompt and instruction conflicts • State and context handling • Memory and retrieval behavior • Stale, contradictory, or incorrectly scoped information • Multi-agent coordination failures • Provenance and source trust • Hallucination containment • Deterministic vs. probabilistic boundaries • Error handling and recovery paths • Retry and fallback behavior • External service failures • Unexpected or adversarial inputs • Observability, logging, and auditability • Hidden assumptions in architecture or workflow design • Cases where individually reasonable components create unreasonable system behavior
A pressure test may include architecture review, structured questioning, adversarial scenarios, counterexamples, boundary cases, controlled failure injection, workflow tracing, and attempts to falsify the assumptions the system depends on.
I am especially interested in questions like:
“What has to remain true for this system to work?”
“What happens when that stops being true?”
“What information can cross this boundary?”
“What happens when two sources disagree?”
“What does the agent do when it is uncertain?”
“What can happen without a human noticing?”
“What state survives a failure?”
“What assumption has nobody tested because everybody already believes it?”
Deliverables depend on scope and may include:
• Findings and prioritized failure modes • Architecture observations • Reproduction steps • Risk and severity notes • Concrete remediation recommendations • Suggested test cases • Improved recovery or permission boundaries • Documentation of unresolved assumptions • Follow-up validation after changes
Relevant VESTIGIA work includes pressure-testing persistent multi-agent systems, memory and continuity architecture, autonomous workflows, source provenance, bounded permissions, human-in-the-loop controls, recovery systems, and research infrastructure.
This is not about generating a long list of hypothetical problems to look impressive.
It is about finding the failures that matter, explaining why they matter, and helping make the system harder to fool — including by its own designers.
If you already have a strange behavior nobody can explain, even better.
Bring the weird bug.
FAQs

Contact for pricing
Duration1 week
Tags
AI Agents
Python
Software Testing
AI Developer
Automation
Generative AI
Quality Assurance
System Architecture
Technical Consulting
Service provided by
Jeff Padget Oxford, USA
AI System Audit & Pressure TestingJeff Padget
Contact for pricing
Duration1 week
Tags
AI Agents
Python
Software Testing
AI Developer
Automation
Generative AI
Quality Assurance
System Architecture
Technical Consulting
Cover image for AI System Audit & Pressure Testing
Your AI system works in the demo.
Now let’s find out what happens when reality stops cooperating.
I audit and pressure-test AI agents, workflows, automations, memory systems, and multi-agent architectures to identify brittle assumptions, hidden failure modes, confusing boundaries, and places where a system can behave correctly according to one component while still failing as a whole.
The goal is not to prove that a system is bad.
The goal is to discover how it can fail while there is still time to do something useful about it.
Depending on the system, testing may examine:
• Agent roles and responsibility boundaries • Tool permissions and autonomy • Human approval checkpoints • Prompt and instruction conflicts • State and context handling • Memory and retrieval behavior • Stale, contradictory, or incorrectly scoped information • Multi-agent coordination failures • Provenance and source trust • Hallucination containment • Deterministic vs. probabilistic boundaries • Error handling and recovery paths • Retry and fallback behavior • External service failures • Unexpected or adversarial inputs • Observability, logging, and auditability • Hidden assumptions in architecture or workflow design • Cases where individually reasonable components create unreasonable system behavior
A pressure test may include architecture review, structured questioning, adversarial scenarios, counterexamples, boundary cases, controlled failure injection, workflow tracing, and attempts to falsify the assumptions the system depends on.
I am especially interested in questions like:
“What has to remain true for this system to work?”
“What happens when that stops being true?”
“What information can cross this boundary?”
“What happens when two sources disagree?”
“What does the agent do when it is uncertain?”
“What can happen without a human noticing?”
“What state survives a failure?”
“What assumption has nobody tested because everybody already believes it?”
Deliverables depend on scope and may include:
• Findings and prioritized failure modes • Architecture observations • Reproduction steps • Risk and severity notes • Concrete remediation recommendations • Suggested test cases • Improved recovery or permission boundaries • Documentation of unresolved assumptions • Follow-up validation after changes
Relevant VESTIGIA work includes pressure-testing persistent multi-agent systems, memory and continuity architecture, autonomous workflows, source provenance, bounded permissions, human-in-the-loop controls, recovery systems, and research infrastructure.
This is not about generating a long list of hypothetical problems to look impressive.
It is about finding the failures that matter, explaining why they matter, and helping make the system harder to fool — including by its own designers.
If you already have a strange behavior nobody can explain, even better.
Bring the weird bug.
FAQs

Contact for pricing