
This project studies how to make the reasoning produced by medical LLMs more factually reliable, not just their final answers. We construct controlled incorrect rationales with known clinically important errors and use them to train and evaluate preference-based models across several model families. Our results suggest that the quality and structure of the negative reasoning examples matter far more than the precise weighting strategy used during training. While multiple-choice accuracy changes little, the models become much better at distinguishing correct from subtly incorrect medical reasoning, highlighting the limits of accuracy alone as a measure of factual robustness.
Team: Sami Rashid, Sumaiya Karim Katha, Ezharuddin Jubaer, Md Fahim, AKM Mahbubur Rahman


