AI Agent System Retainer: evals, monitoring, incident response by Alikhan UrumovAI Agent System Retainer: evals, monitoring, incident response by Alikhan Urumov
AI Agent System Retainer: evals, monitoring, incident responseAlikhan Urumov
Cover image for AI Agent System Retainer: evals, monitoring, incident response

Who this is for

A system that is already live, already spending money, and already making decisions your business depends on. You do not have to have built it with me, but I will run an audit before I take standing responsibility for it.
This is the wrong product for a system that is not yet carrying real traffic, and I will tell you that rather than sell it to you.

What you are buying

A standard that stays true. Your system runs against written thresholds — task success rate against your eval set, p95 latency, cost per task, uptime — and my job is that it keeps meeting them while the ground moves underneath it. Traffic changes. Providers deprecate models and reprice them. Your own upstream systems change. The thresholds do not.

The boundary

One deployed system: the agents, services and integrations named on a covered-systems list we agree before the first month starts. That list is the retainer. Anything not on it is a build.

Inside the retainer

Continuous monitoring against the agreed thresholds, with alerting tuned so that an alert means something
Scheduled eval and regression runs against live traffic rather than a fixture set, with the results sent to you whether or not they are good
Cost held inside its threshold, with the reductions actioned rather than reported and left
Model, prompt and dependency changes tested against the eval suite before they ship, not after
Incident response under the commitment below
One scheduled improvement each month — a capability, agent or integration inside the covered system, in your priority order — delivered against its own written criteria and its own eval, exactly the way a build is
A written operations report each month: what ran, what moved, what broke, what it cost, what changed, whether every threshold still holds, and what I recommend next. Short, dated, and the thing you hand to your board or your CTO
Every month the system is better tested and better documented than it was at the start of it. The eval suite grows: every real failure becomes a permanent test case. The runbook grows: every incident closes with the entry that would have made it boring. At least one new decision record is added, so next year's engineer knows why this month's change was made.
All of it lives in your repository and your cloud account. None of it stops being useful if you stop paying me.

Outside the retainer, and quoted as a build

A new workflow, or an agent set that is not on the covered-systems list
An integration with a system that is not on that list
Replatforming, cloud migration, or a rewrite
Anything with its own acceptance criteria. It gets its own written scope and its own fixed price
A second system. Each one goes on the covered-systems list at its own price
Nothing outside the boundary is billed to you after the fact. It is scoped, quoted in writing, and you decide.

If a threshold is breached

Restoring it is inside the retainer. It is not a change request and it does not generate an invoice. That is the whole reason this costs what it costs: a regression is my problem.

When the thresholds stop describing reality

Every report says whether the numbers still describe the system you actually need. When they stop — ten times the traffic, a changed workflow, a new constraint on you — we re-baseline in writing before the next month. Never retroactively, and never as a way of quietly moving a target I have missed.

Response commitment

Critical, meaning the system is down or producing wrong output in production: acknowledged within four working hours and worked continuously until it is stable. Everything else: next business day. My working week is Monday to Friday, 09:00-18:00 UTC+4, which covers European business hours in full and the start of the US East Coast day.
This is a business-hours retainer, not 24/7 on-call. If you need overnight cover, say so before we start and it gets priced properly rather than pretended.

Terms

Month to month, thirty days' notice either way, no exit fee. Everything stays in your repository and your cloud account, so ending the retainer changes nothing about who can operate the system.

The price

$2,500 a month buys standing accountability for a live system's reliability, cost and quality, and the written record that proves it. It is priced against what your system spends and what it costs you when it is wrong — not against the size of the build that produced it.
To start: message me with what is running now and what woke you up last month.
FAQs

Starting at$2,500 /mo
Tags
AI Developer
Backend Engineer
Service provided by
Alikhan Urumov Lelystad, Netherlands
AI Agent System Retainer: evals, monitoring, incident responseAlikhan Urumov
Starting at$2,500 /mo
Tags
AI Developer
Backend Engineer
Cover image for AI Agent System Retainer: evals, monitoring, incident response

Who this is for

A system that is already live, already spending money, and already making decisions your business depends on. You do not have to have built it with me, but I will run an audit before I take standing responsibility for it.
This is the wrong product for a system that is not yet carrying real traffic, and I will tell you that rather than sell it to you.

What you are buying

A standard that stays true. Your system runs against written thresholds — task success rate against your eval set, p95 latency, cost per task, uptime — and my job is that it keeps meeting them while the ground moves underneath it. Traffic changes. Providers deprecate models and reprice them. Your own upstream systems change. The thresholds do not.

The boundary

One deployed system: the agents, services and integrations named on a covered-systems list we agree before the first month starts. That list is the retainer. Anything not on it is a build.

Inside the retainer

Continuous monitoring against the agreed thresholds, with alerting tuned so that an alert means something
Scheduled eval and regression runs against live traffic rather than a fixture set, with the results sent to you whether or not they are good
Cost held inside its threshold, with the reductions actioned rather than reported and left
Model, prompt and dependency changes tested against the eval suite before they ship, not after
Incident response under the commitment below
One scheduled improvement each month — a capability, agent or integration inside the covered system, in your priority order — delivered against its own written criteria and its own eval, exactly the way a build is
A written operations report each month: what ran, what moved, what broke, what it cost, what changed, whether every threshold still holds, and what I recommend next. Short, dated, and the thing you hand to your board or your CTO
Every month the system is better tested and better documented than it was at the start of it. The eval suite grows: every real failure becomes a permanent test case. The runbook grows: every incident closes with the entry that would have made it boring. At least one new decision record is added, so next year's engineer knows why this month's change was made.
All of it lives in your repository and your cloud account. None of it stops being useful if you stop paying me.

Outside the retainer, and quoted as a build

A new workflow, or an agent set that is not on the covered-systems list
An integration with a system that is not on that list
Replatforming, cloud migration, or a rewrite
Anything with its own acceptance criteria. It gets its own written scope and its own fixed price
A second system. Each one goes on the covered-systems list at its own price
Nothing outside the boundary is billed to you after the fact. It is scoped, quoted in writing, and you decide.

If a threshold is breached

Restoring it is inside the retainer. It is not a change request and it does not generate an invoice. That is the whole reason this costs what it costs: a regression is my problem.

When the thresholds stop describing reality

Every report says whether the numbers still describe the system you actually need. When they stop — ten times the traffic, a changed workflow, a new constraint on you — we re-baseline in writing before the next month. Never retroactively, and never as a way of quietly moving a target I have missed.

Response commitment

Critical, meaning the system is down or producing wrong output in production: acknowledged within four working hours and worked continuously until it is stable. Everything else: next business day. My working week is Monday to Friday, 09:00-18:00 UTC+4, which covers European business hours in full and the start of the US East Coast day.
This is a business-hours retainer, not 24/7 on-call. If you need overnight cover, say so before we start and it gets priced properly rather than pretended.

Terms

Month to month, thirty days' notice either way, no exit fee. Everything stays in your repository and your cloud account, so ending the retainer changes nothing about who can operate the system.

The price

$2,500 a month buys standing accountability for a live system's reliability, cost and quality, and the written record that proves it. It is priced against what your system spends and what it costs you when it is wrong — not against the size of the build that produced it.
To start: message me with what is running now and what woke you up last month.
FAQs

$2,500 /mo