I lead hands-on Rails and PostgreSQL delivery for a production SaaS platform. In one verified incident, I traced a three-hour outage with green process health to stale PostgreSQL connections consuming request capacity, introduced bounded client-side failure detection, validated the change in staging and production, and added a dedicated 504 alert. The same monitored 504 pattern was not observed again during the following month.