Counterpoint: I've seen rubric-first make models overconfident on edge cases. Did you check calibration?