AI Agent System Audit: cost, failure modes, written criteria by Alikhan UrumovAI Agent System Audit: cost, failure modes, written criteria by Alikhan Urumov
AI Agent System Audit: cost, failure modes, written criteriaAlikhan Urumov
Cover image for AI Agent System Audit: cost, failure modes, written criteria

What this is

A review of one AI or agent system you already run in production, ending in a document that states what that system has to do in numbers — and what it currently does instead.
You are not buying an opinion. Every finding carries a reproduction: the command, the query or the trace that shows it on your own system.

The boundary

One system. One repository, one deployed environment, and the integrations that system already calls. That is the whole scope.
The boundary is fixed, which is what makes the price fixed. What I find changes the report; it does not change what you pay.

Inside the boundary

Architecture: agent boundaries, orchestration, state and memory, the tool and API surface
Reliability: failure modes, retries, timeouts, idempotency, and where a human belongs in the loop
Cost: spend per task split across models, retrieval, retries and infrastructure, taken from your own billing data rather than estimated
Quality: whether task success is measured at all, and what the number is once it is
Observability: what is traced today, what is not, and the evals worth having first
Security: prompt-injection surface, secret handling, data egress

Outside the boundary

A second system, or a second deployment of the same one. Each is its own audit
Writing or fixing code. This ends in a plan, not a pull request
Model training, fine-tuning and dataset work
Frontend, design and product strategy
Anything not yet in production. If it does not run, there is nothing to measure
If your scope genuinely has to widen, you get a quote before anything further is read, and you are free to say no.

What you end up owning

Six documents, delivered as files you keep, written to be handed to someone else — including an engineer who is not me.
An acceptance criteria document: task success rate against a named eval set, p95 latency, cost per task, uptime target, and the required behaviour when a model or an upstream API misbehaves. Written as targets any engineer can be held to, me or anyone else.
A measured baseline showing where your system sits against each of those targets today, taken from your own infrastructure and your own bills. Nothing in it is a benchmark I ran somewhere else.
An architecture diagram of the system as it actually runs, not as the docs describe it. Delivered as an editable source file, not a screenshot.
A findings register, ranked by what each issue costs you rather than by technical severity, every entry carrying its reproduction.
A prioritised fix plan in which every item states the criterion it has to satisfy to count as done. It is written as the scope document for whoever builds next, so you are not paying someone else to write one.
A recorded walkthrough of the findings, so the reasoning survives the people who happened to be in the room.

The standard this work is held to

Every claim in the report reproduces on your system from the instructions inside it. If one does not, say so: I either reproduce it with you, or I withdraw it and reissue the fix plan without it at no charge.

What happens to the fee

If the finding is that your agent system is not your problem — that it is a data problem, a product problem, or something that does not need building at all — the report says exactly that. A diagnosis that can only ever recommend buying the next thing is not a diagnosis.
The fee is credited in full against a build booked within 30 days of delivery. If you go ahead, the audit has cost you nothing.
You keep all six documents in every case.

What I need from you

Read access to the repository, read access to your cloud and observability dashboards, and any existing docs or diagrams. If nothing is traced, that is a finding, not a blocker.
To start: message me with what is running, what keeps breaking, and read access. Questions get answered in writing the same working day. No call needed unless you want one.
FAQs

Starting at$4,000
Duration2 weeks
Tags
AI Developer
Backend Engineer
Service provided by
Alikhan Urumov Lelystad, Netherlands
AI Agent System Audit: cost, failure modes, written criteriaAlikhan Urumov
Starting at$4,000
Duration2 weeks
Tags
AI Developer
Backend Engineer
Cover image for AI Agent System Audit: cost, failure modes, written criteria

What this is

A review of one AI or agent system you already run in production, ending in a document that states what that system has to do in numbers — and what it currently does instead.
You are not buying an opinion. Every finding carries a reproduction: the command, the query or the trace that shows it on your own system.

The boundary

One system. One repository, one deployed environment, and the integrations that system already calls. That is the whole scope.
The boundary is fixed, which is what makes the price fixed. What I find changes the report; it does not change what you pay.

Inside the boundary

Architecture: agent boundaries, orchestration, state and memory, the tool and API surface
Reliability: failure modes, retries, timeouts, idempotency, and where a human belongs in the loop
Cost: spend per task split across models, retrieval, retries and infrastructure, taken from your own billing data rather than estimated
Quality: whether task success is measured at all, and what the number is once it is
Observability: what is traced today, what is not, and the evals worth having first
Security: prompt-injection surface, secret handling, data egress

Outside the boundary

A second system, or a second deployment of the same one. Each is its own audit
Writing or fixing code. This ends in a plan, not a pull request
Model training, fine-tuning and dataset work
Frontend, design and product strategy
Anything not yet in production. If it does not run, there is nothing to measure
If your scope genuinely has to widen, you get a quote before anything further is read, and you are free to say no.

What you end up owning

Six documents, delivered as files you keep, written to be handed to someone else — including an engineer who is not me.
An acceptance criteria document: task success rate against a named eval set, p95 latency, cost per task, uptime target, and the required behaviour when a model or an upstream API misbehaves. Written as targets any engineer can be held to, me or anyone else.
A measured baseline showing where your system sits against each of those targets today, taken from your own infrastructure and your own bills. Nothing in it is a benchmark I ran somewhere else.
An architecture diagram of the system as it actually runs, not as the docs describe it. Delivered as an editable source file, not a screenshot.
A findings register, ranked by what each issue costs you rather than by technical severity, every entry carrying its reproduction.
A prioritised fix plan in which every item states the criterion it has to satisfy to count as done. It is written as the scope document for whoever builds next, so you are not paying someone else to write one.
A recorded walkthrough of the findings, so the reasoning survives the people who happened to be in the room.

The standard this work is held to

Every claim in the report reproduces on your system from the instructions inside it. If one does not, say so: I either reproduce it with you, or I withdraw it and reissue the fix plan without it at no charge.

What happens to the fee

If the finding is that your agent system is not your problem — that it is a data problem, a product problem, or something that does not need building at all — the report says exactly that. A diagnosis that can only ever recommend buying the next thing is not a diagnosis.
The fee is credited in full against a build booked within 30 days of delivery. If you go ahead, the audit has cost you nothing.
You keep all six documents in every case.

What I need from you

Read access to the repository, read access to your cloud and observability dashboards, and any existing docs or diagrams. If nothing is traced, that is a finding, not a blocker.
To start: message me with what is running, what keeps breaking, and read access. Questions get answered in writing the same working day. No call needed unless you want one.
FAQs

$4,000