You have two models. Model A has F1 of 0.82. Model B has F1 of 0.79.

Model A wins, right?

Not necessarily. F1 is calculated at one specific threshold. Maybe Model B is much better at other thresholds. Maybe on your actual deployment threshold, B beats A.

ROC curves show you the full picture. They plot model performance across every possible...