⚖️ 2. Sizing via Empirical p99 Telemetry: Instead of arbitrary downsizing, we ingested historical p99 CPU, memory, and disk IOPS metrics across all machine families. Production nodes had been oversized to accommodate traffic spikes that only occurred two hours a day. We rightsized instance families, introduced ephemeral scheduling for non-production environments, and matched compute envelopes to empirical demand.