Modern AI applications can fail in ways that traditional software testing does not easily catch. A prompt change can reduce answer quality, a model upgrade can introduce regressions, an agent can expose sensitive information, or a jailbreak can bypass application-level safeguards. Arbiter provides an infrastructure layer for continuously evaluating and controlling these systems.