Beyond model training, I handled the GPU infrastructure end-to-end: VRAM sizing, AWS EC2 provisioning, quota management, spot-instance fallback, cost-controlled start/stop workflows, and debugging failed training runs with smoke tests before full experiments.