I build the training set, fine-tune an open model with QLoRA, and evaluate it against your current baseline on held-out data. That includes a check that the model has not gotten worse at everything else, which is the failure most fine-tuning work misses.