Logrine
Synthetic data Evaluation Blog About Trust Careers Book a scoping call
← All articles
Evaluation 11 October 2026 · 7 min read

LLM-as-a-judge: three biases to control before you trust a score

Scoring open-ended model output by hand does not scale, so most teams now use a language model as the judge. It is a practical approach, and it is also easy to misuse. A judge score is a measurement, and a measurement needs to be checked before it is believed.

What the research shows

Zheng and colleagues studied this in "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (NeurIPS 2023, Datasets and Benchmarks track). They found that a strong judge model can reach over 80% agreement with human preferences, which is about the level at which humans agree with each other. They also documented systematic weaknesses. Agreement on average does not mean the judge is unbiased on any given dataset, and that is the point of calibrating it.

Three biases to control

1. Position bias

When a judge compares two answers, it can favour whichever one appears first (or second), regardless of quality. Control: run every comparison twice with the order swapped, and count a verdict only when both runs agree. Treat inconsistent results as ties, and report how often they occur.

2. Verbosity bias

Judges can prefer longer answers even when the extra length adds nothing. This matters most when you are comparing models that differ in how wordy they are. Control: score against explicit criteria rather than an overall impression, penalize padding in the rubric, and check whether your scores correlate with answer length.

3. Self-preference bias

The same paper discusses a tendency for a judge to favour answers written by itself or by similar models, with evidence that is suggestive rather than conclusive. The risk is practical: if the same model family both generates and judges, errors can go unnoticed. Control: use a judge from a different model family than the generator, and spot-check with human review.

A related limit: grading reasoning

The same study notes that judges are less reliable when grading math and reasoning answers. Control: where a verified answer exists, give it to the judge as a reference, or check the final answer programmatically instead of by judgment.

A calibration routine

  1. Write a rubric with explicit criteria and a description of each score level.
  2. Have humans label a sample, stratified by the slices you care about. Double-label part of it to measure how much humans agree with each other, which is the realistic ceiling for the judge.
  3. Run the judge on the same sample.
  4. Measure agreement, including a chance-corrected statistic such as Cohen's kappa, not just raw percent agreement.
  5. Read the disagreements by slice. Revise the rubric where the judge and humans diverge for a reason you can name.
  6. Re-run, then freeze the judge model version and prompt used for scoring.
  7. Repeat the check whenever the judge, the rubric or the task changes.

What a trustworthy report states

The takeaway

LLM judges are a useful tool, not an oracle. Calibrate against humans, test for the known biases, and publish the evidence next to the score. A benchmark result you cannot audit is an opinion.

Sources
  • Zheng, L., Chiang, W.-L., Sheng, Y., et al. "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." NeurIPS 2023, Datasets and Benchmarks Track.
  • Cohen, J. "A coefficient of agreement for nominal scales." Educational and Psychological Measurement, 1960.
Keep reading
Model collapse: why synthetic data needs a real-data anchor
What belongs in a dataset card for synthetic data
Logrine

Synthetic datasets and evaluation benchmarks for teams shipping LLMs and AI agents, delivered with the evidence.

Product
Synthetic data Evaluation benchmarks Trust and data handling
Company
About Careers Contact hello@logrine.com
Resources
Blog Model collapse LLM-as-a-judge Dataset cards
© 2026 MaxxLabs. Logrine is a MaxxLabs product.
Bengaluru, India