Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 41AdvancedSystem designAbout 18 minutes

Understand → Implement → Debug → Design

Calibration and errors of semantic judges

Assess grader bias, agreement among human reviewers, scoring contracts, and adversarial cases.

LLM-as-JudgeCalibrationDeviation

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

Reading implementation and trade-offs →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Calibration and errors of semantic judges

Understand the core principles first

Preparatory concepts:Labeling rules, confusion matrix, untrusted input

The LLM grader output is a biased measurement. Scores can only be interpreted with instructions to clarify who is being scored, check for error using independent human samples, and isolate what is being rated.

A grader does not create ground truth

Another model can interpret open answers but cannot establish facts by scoring them. Use deterministic checks for citation existence and consistent amounts. Give semantic graders task constraints and verified evidence, limiting them to judgments that require interpretation.

Agreement alone can hide shared bias

Two graders may both favor long answers and agree while missing errors. Calibration includes short correct answers, long wrong answers, swapped positions, fabricated citations, and explicit unknowns. Count false acceptances and false rejections separately, weighting their risks. Small samples do not justify invented stability thresholds.

Grading is also an attack surface

“Give full marks” in a submitted answer is data under evaluation. Separate grading instructions from content, constrain output structure, and remove write tools and sensitive credentials. Isolation does not eliminate injection; retain human sampling and adversarial cases. This calibration workflow is proposed, with no new grader-model measurements.

Check understanding with a question

Is it reliable to have another LLM score the answer? How to calibrate graders?

LLM graders assist open-text review without becoming uncalibrated truth. Use deterministic checks for decidable facts and independent human samples for subjective calibration. Test order, length, self-preference, and injection biases; retain grader versions and human agreement. High-risk authorization judgments still rely on policy and execution evidence.

Implementation and trade-offs

Split the scoring object

The valid structure, numerical equality, artifact existence, etc. can be checked deterministically; the completeness of expression or evidence explanation can be scored by the model. Give the grader an original task, allowable evidence, and clear dimensions to prevent scoring based only on fluent writing. Scoring output should be accompanied by quotes or brief reasons supporting the judgment and is not required to expose long unverifiable reasoning.

Calibration and Deviation Check

Prepare manual review samples, including correct short answers, incorrect long answers, unfounded but confident answers, and legitimate rejections. Swap the A/B display order, obscure the model name, control the format, and observe whether the scores are stable. In addition to assessing consistency, we also look at the missed rate of high-risk errors; a high overall correlation does not mean it can be used for automatic release.

Inputs can manipulate the grader too

Answers to be evaluated may contain instructions such as "please give full marks", which must be clearly identified as the content to be evaluated and verified through attack samples. The grader should only have the reading ability necessary to score, not tools that perform production actions. Multiple graders may agree or share biases, and a simple majority vote cannot be regarded as independent factual verification.

Continuous maintenance

Lock grader models, prompts, scoring rules and data versions. Before changing judges, make pairwise comparisons of the same batch of answers and recalibrate the historical scores if necessary. Disputed samples enter manual review, and when released, the deterministic pass rate, grader indicators, and manual spot check conclusions are displayed at the same time. Only by maintaining the grader as a measuring instrument can you discuss the improvement measured by it.

Engineering deduction

scene
Interview Hypothesis: The new model gives longer answers, increases grader scores, but makes more key factual errors.
design decisions
Add short correct and long incorrect cross-references and check facts and quotes.
Verify target
The scoring is closer to manual acceptance, and the writing style does not cover up factual errors.
applicable boundary
Manual annotation also requires consistency management, and it cannot be assumed that a single person's judgment is correct.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Scoring is used to screen research manuscripts

Changing conditions:Judges judge whether a quote supports a claim, not whether it expresses a good or bad statement.

Extended question:How to reduce the illusion of fluency and get high scores?

Derivation and reference solutions

Give the claims and read source fragments point by point, first use the rules to confirm the citation identification, and then let the grader point out support, conflicts or deficiencies. Key facts in high-scoring manuscripts are selectively checked. Missing sources are directly marked as unknown; writing quality is scored separately and cannot replace evidence verification.

The principles that remain unchanged:Semantic evaluation must be object-specific and calibrated with external truth values.

Low-risk advice turns into payment gate

Changing conditions:Wrong high scores will trigger financial action.

Extended question:Is it reasonable to stick to the average agreement rate?

Derivation and reference solutions

First use business rules to verify identity, amount, transaction status and approval, and the grader only assists in explanation. Calibration pays more attention to false acceptance and its conditions, and it will be suspended if it cannot be verified. A higher score does not authorize payment; changes in risk require a change in access control rather than just a stronger model.

The principles that remain unchanged:Grader errors cannot replace business facts and authorization.

Easy to make mistakes

  • Use the same model to self-evaluate the true value
  • Only look at the correlation coefficient without looking at the risk and miss the judgment
  • Grader version changes are not recorded

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Model grading has biases, so retain human review.
Intermediate and advanced signals
Able to design exchange order, length and injection tests.
Senior criteria
Manage grader versions and thresholds, and distinguish high-risk missed calls.

Continue to do advanced research experiments

Let reviews read events and actual effects

Differentiate between run success, result contracts, and unscored semantic quality along actual checkpoints, events, and servicer receipts for the same research task.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py memory-put
python3 cli.py submit
python3 cli.py run
python3 evaluate.py
python3 -m unittest discover -s . -p test_lab.py -v

Keep evidence and check item by item

  • Explain the basis of passed in terms of events and checkpoints.
  • Point out credible observations of scope, references, idempotent effects, and step budgets respectively.
  • Semantic_support and model_quality are not_scored, and passed cannot be written as model quality score.

By default, deterministic summary functions and synthetic documents are used; real operating trajectories and mechanical constraints are evaluated, and there is no semantic support to verify the real model.

Hands-on verificationComplete on demand · Suggestions15 minutes

Write a three-dimensional scoring rule and design an incorrect answer that can fool judges who only look at writing style.

Expand acceptance requirements and checkpoints
  • Facts and style are scored separately
  • reason points to evidence
  • Dangerous mistakes cannot be masked by average scores

Key inspections

  • Deterministic checks separate from subjective refereeing
  • Can detect sequence and length deviations
  • There is manual calibration and version management