A benchmark is a signal. Read the context.
Model scores become useful when you understand the task, evaluation method, and tradeoffs behind them.

A leaderboard can help you build a shortlist. It cannot tell you how a model will behave in every product. The useful question is whether the evaluation resembles the work you need done.
Read beyond the position
Check the task, dataset, evaluation date, model version, and prompting setup. A score from one test is not interchangeable with a score from another. An unreported result means the evidence is missing; it does not mean the model scored zero.
Compare more than accuracy
Stanford's HELM research argues for evaluating multiple dimensions and recognizing gaps in coverage. For a product, consider reliability, response time, cost, and behavior on difficult inputs alongside task performance.
Build a small evaluation of your own
Collect representative examples and decide what a successful answer looks like before trying models. Include edge cases, record failures, and keep the setup consistent. Recheck when the model or prompt changes.