OPEN MEDICAL AI BENCHMARKS / BY MEDARC × SOPHONT

Medmarks Leaderboard

Paper Code Blog Post Join us on Discord
How to interpret win rates

A win rate summarizes how well a model performs relative to the other models in the evaluation. For each question, a model earns 1 point for scoring higher than another model, 0.5 points for tying, and 0 points for scoring lower. The final win rate combines these comparisons into a weighted average across datasets and models. A score of 64% means the model earns 64% of the possible comparison points against its peers, not that it answers 64% of questions correctly. See the paper’s scoring method.

As a rule of thumb, win rates less than 0.5 percentage points apart are too close to meaningfully distinguish. A gap of 0.5 to less than 1 point is a small lead, 1 to less than 3 points is a notable lead, and 3 points or more is a substantial lead. For example, 64.0% versus 64.3% is a close result, while 66.0% represents a notable two-point lead over 64.0%.

News

Acknowledgements

We are grateful to Prime Intellect for their generous support in running proprietary model APIs through their Inference platform. Thanks to FAL AI for providing a compute grant that helped support this research.

If you are a model developer/frontier lab, we'd love to have your model added to our leaderboard. Please contact us!