LLM as a judge: useful reviewer, imperfect measurement
Design model-based grading with explicit rubrics, position checks, human calibration, and deterministic checks where they fit.

A model can help review large collections of generated answers, but a second model's opinion does not turn an answer into a verified fact. The grading system has its own prompt, assumptions, and failure modes. Evaluate it with the same care you apply to the system it is grading.
The MT-Bench and Chatbot Arena research investigated model judges and discussed biases including response position and verbosity. These findings are a reason to inspect grading behavior, not a universal claim that every judge fails in the same way.
Define a rubric that describes the task
Suppose an assistant writes product-support replies. A useful rubric separates factual support, completeness, tone, and adherence to the requested format. “Which answer is better?” combines these dimensions without stating their relative importance.
Include examples of unacceptable errors. A courteous answer that invents a refund policy should fail the factual-support criterion. A brief answer that correctly asks for a missing order number should not lose simply because another answer is longer. State those distinctions before the judge sees the candidate responses.
Use code where code can decide
Check parseability, required fields, exact identifiers, and executable test results with deterministic validators when appropriate. A model judge adds little value to deciding whether an object contains a required key. It can help with qualities that are harder to encode, such as whether an explanation addresses the user's actual concern.
Keep the outputs of different checks separate. A response may pass the format validator and fail the factual review. A single blended number makes it harder to identify what needs fixing and can let a strong style score hide a serious error.
Calibrate against reviewed examples
Create a set of answers that human reviewers have assessed with the same rubric. Resolve disagreements about the rubric itself before treating that set as a reference. Then inspect where the model judge disagrees and whether those disagreements cluster around particular error types.
Use the result to decide the judge's role. It might be suitable for prioritizing a review queue while remaining unsuitable as an automatic release gate. The acceptable level of disagreement depends on the consequences of a wrong acceptance, not just on the volume of answers you want to process.
Probe position and style effects
For pairwise comparisons, swap the order of the answers and check whether the decision changes. Compare concise and verbose versions that contain the same relevant facts. Keep model identity hidden when it is not part of the task, and avoid unnecessary metadata that could influence preference.
Treat candidate answers as untrusted content. A response may contain text telling the judge to award a high score. The grading prompt should identify that text as material to assess, and the surrounding workflow should include tests for such attempts. Prompt wording alone is not proof that the judge is immune.
Report uncertainty in the measurement
Version the judge model, prompt, rubric, and sampling settings. Inspect disagreements and borderline cases rather than reporting only an overall pass rate. Reassess the grading setup when changing the kinds of answers being judged.
Use model grading to extend a review process that has clear standards and reference evidence. Keep periodic human review and deterministic checks in place. An efficient grading pipeline is valuable when its limitations are visible enough to guide decisions.