I built a production-shaped reliability platform for ephemeral GitHub Actions runner fleets. It turns queue, startup, cleanup, capacity, incident, and cost signals into repeatable operational evidence.
Reliability model
The platform defines and evaluates SLOs for:
Queue latency
Runner startup success
Cleanup completion
Maximum job age
Prometheus metrics and alert rules expose reliability signals, while the API stores evidence and produces capacity and incident recommendations.
Failure testing
Deterministic scenarios exercise:
Workload bursts
Runner capacity loss
Image-pull failures
GitHub API degradation
Each scenario produces evidence showing which SLOs pass or fail rather than relying on a polished demonstration alone.
Infrastructure and delivery
Python and FastAPI control API
OpenTofu AWS EKS lab design
Kubernetes and Helm deployment
Prometheus metrics and alerts
GitHub Actions CI
Docker packaging
Read-only MCP tools for fleet inspection and recommendations
Public CI validates Python, MCP integration, OpenTofu, containers, Helm, and a disposable Kubernetes deployment.
Safety and evidence boundaries
The repository uses public and synthetic material. The AWS environment is a guarded temporary-lab design with budget and teardown controls. Production cloud, GPU, and paying-customer usage are not claimed.