Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
Practice 41AdvancedSystem designAbout 18 minutes

Corresponding knowledge: Calibration and errors of semantic judges

Is it reliable to have another LLM score the answer? How to calibrate graders?

Assess grader bias, agreement among human reviewers, scoring contracts, and adversarial cases.

LLM-as-JudgeCalibrationDeviation

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge exercises·Independent answers

My notes and review ↗

Principles and Solutions have been collapsed. Explain the core mechanism, boundaries and verification methods in your own words, and then compare them.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Explain in your own words first

The core principles, analysis, Q&A and migration cases have been closed. When you are ready, unfold it and compare it with the content to find any omissions.

Hands-on verificationComplete on demand · Suggestions15 minutes

Write a three-dimensional scoring rule and design an incorrect answer that can fool judges who only look at writing style.

Expand acceptance requirements and checkpoints
  • Facts and style are scored separately
  • reason points to evidence
  • Dangerous mistakes cannot be masked by average scores

Key inspections

  • Deterministic checks separate from subjective refereeing
  • Can detect sequence and length deviations
  • There is manual calibration and version management