Runner Fleet Reliability Platform by Joseph JosephRunner Fleet Reliability Platform by Joseph Joseph

Runner Fleet Reliability Platform

Joseph Joseph

Joseph Joseph

Overview

I built a production-shaped reliability platform for ephemeral GitHub Actions runner fleets. It turns queue, startup, cleanup, capacity, incident, and cost signals into repeatable operational evidence.

Reliability model

The platform defines and evaluates SLOs for:
Queue latency
Runner startup success
Cleanup completion
Maximum job age
Prometheus metrics and alert rules expose reliability signals, while the API stores evidence and produces capacity and incident recommendations.

Failure testing

Deterministic scenarios exercise:
Workload bursts
Runner capacity loss
Image-pull failures
GitHub API degradation
Each scenario produces evidence showing which SLOs pass or fail rather than relying on a polished demonstration alone.

Infrastructure and delivery

Python and FastAPI control API
OpenTofu AWS EKS lab design
Kubernetes and Helm deployment
Prometheus metrics and alerts
GitHub Actions CI
Docker packaging
Read-only MCP tools for fleet inspection and recommendations
Public CI validates Python, MCP integration, OpenTofu, containers, Helm, and a disposable Kubernetes deployment.

Safety and evidence boundaries

The repository uses public and synthetic material. The AWS environment is a guarded temporary-lab design with budget and teardown controls. Production cloud, GPU, and paying-customer usage are not claimed.
Like this project

Posted Sep 30, 2026

Production-shaped SRE platform for ephemeral CI/CD runners with SLOs, observability, capacity modeling, failure simulations, IaC, and read-only MCP operations.