Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Assess grader bias, agreement among human reviewers, scoring contracts, and adversarial cases.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
RAG evidence flow and failure diagnosis →Idempotency, unknown outcomes, and task recovery →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
Reading implementation and trade-offs →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Calibration and errors of semantic judges
Preparatory concepts:Labeling rules, confusion matrix, untrusted input
The LLM grader output is a biased measurement. Scores can only be interpreted with instructions to clarify who is being scored, check for error using independent human samples, and isolate what is being rated.
Another model can interpret open answers but cannot establish facts by scoring them. Use deterministic checks for citation existence and consistent amounts. Give semantic graders task constraints and verified evidence, limiting them to judgments that require interpretation.
Two graders may both favor long answers and agree while missing errors. Calibration includes short correct answers, long wrong answers, swapped positions, fabricated citations, and explicit unknowns. Count false acceptances and false rejections separately, weighting their risks. Small samples do not justify invented stability thresholds.
“Give full marks” in a submitted answer is data under evaluation. Separate grading instructions from content, constrain output structure, and remove write tools and sensitive credentials. Isolation does not eliminate injection; retain human sampling and adversarial cases. This calibration workflow is proposed, with no new grader-model measurements.
LLM graders assist open-text review without becoming uncalibrated truth. Use deterministic checks for decidable facts and independent human samples for subjective calibration. Test order, length, self-preference, and injection biases; retain grader versions and human agreement. High-risk authorization judgments still rely on policy and execution evidence.
The valid structure, numerical equality, artifact existence, etc. can be checked deterministically; the completeness of expression or evidence explanation can be scored by the model. Give the grader an original task, allowable evidence, and clear dimensions to prevent scoring based only on fluent writing. Scoring output should be accompanied by quotes or brief reasons supporting the judgment and is not required to expose long unverifiable reasoning.
Prepare manual review samples, including correct short answers, incorrect long answers, unfounded but confident answers, and legitimate rejections. Swap the A/B display order, obscure the model name, control the format, and observe whether the scores are stable. In addition to assessing consistency, we also look at the missed rate of high-risk errors; a high overall correlation does not mean it can be used for automatic release.
Answers to be evaluated may contain instructions such as "please give full marks", which must be clearly identified as the content to be evaluated and verified through attack samples. The grader should only have the reading ability necessary to score, not tools that perform production actions. Multiple graders may agree or share biases, and a simple majority vote cannot be regarded as independent factual verification.
Lock grader models, prompts, scoring rules and data versions. Before changing judges, make pairwise comparisons of the same batch of answers and recalibrate the historical scores if necessary. Disputed samples enter manual review, and when released, the deterministic pass rate, grader indicators, and manual spot check conclusions are displayed at the same time. Only by maintaining the grader as a measuring instrument can you discuss the improvement measured by it.
Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1Are the two graders necessarily correct if they agree?
The main question talks about reliability and continues to check whether the consensus produces independent evidence.
Not necessarily. Identical training preferences, similar cues, and the same evidence gaps produce correlated errors. Use difficult samples marked by independent experts to evaluate the false-acceptance rate and check the types of errors shared by both parties; multiple graders can indicate uncertainty, but unanimous opinions cannot replace fact sources and hard constraint checks.
Level 1What should I do if there are grading instructions hidden in the answers?
After the judges read the answers, the answers themselves may attempt to control the scoring.
Treat answers as non-command input with boundaries, put the scoring rules in the trusted layer, limit the grader's ability and output, and add scoring injection counterexamples. Use programmatic checks for properties that can be verified deterministically; if the model changes its score due to "ignoring the rules", the grader will be recorded as invalid and transferred to manual work. Just filtering a certain keyword cannot cover encoding, citations, and multiple rounds of attacks.
Follow this answer further
Level 2Does wrapping the answer into a JSON string completely solve the problem of score injection?
The parent question uses input isolation, and the child question checks whether the serialization is equal to semantic security.
No. JSON guarantees structural escaping, not that the model will ignore instructions inside a string. Structured packaging helps distinguish fields, but the judge still needs a narrowly defined task and no external execution capabilities. Test semantic injection, and use rule checks and human review for critical facts.
Follow this answer further
Level 3Graders only output scores and have no tools. Will there be any actual risk if the scores are manipulated?
The father asked about the risk of shrinking grader authority and continuing to pursue scoring as a downstream control signal.
Yes. False high scores feed into downstream decisions if the score determines automatic releases, refunds, or model upgrades. Allow high-risk decisions to be accompanied by factual gates, permission checks and approval; treat graders as one of the evidence and retain reviewable reasons. Just because the grader has no tools does not mean that the entire system has no side effects.
Level 1Can historical scores still be compared directly after changing graders?
Calibration is bound to the grader version, and the comparative definition needs to be restored after upgrading.
Absolute scores cannot be compared directly. Freeze a set of manually labeled anchor points so that old and new graders can re-evaluate under the same conditions and observe ordering, false acceptance and scale differences; record models, prompts, rules and evidence versions. The two sets of historical curves can be reported separately without applying arbitrary linear scaling to obscure classification differences.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:Judges judge whether a quote supports a claim, not whether it expresses a good or bad statement.
Extended question:How to reduce the illusion of fluency and get high scores?
Give the claims and read source fragments point by point, first use the rules to confirm the citation identification, and then let the grader point out support, conflicts or deficiencies. Key facts in high-scoring manuscripts are selectively checked. Missing sources are directly marked as unknown; writing quality is scored separately and cannot replace evidence verification.
The principles that remain unchanged:Semantic evaluation must be object-specific and calibrated with external truth values.
Changing conditions:Wrong high scores will trigger financial action.
Extended question:Is it reasonable to stick to the average agreement rate?
First use business rules to verify identity, amount, transaction status and approval, and the grader only assists in explanation. Calibration pays more attention to false acceptance and its conditions, and it will be suspended if it cannot be verified. A higher score does not authorize payment; changes in risk require a change in access control rather than just a stronger model.
The principles that remain unchanged:Grader errors cannot replace business facts and authorization.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
Differentiate between run success, result contracts, and unscored semantic quality along actual checkpoints, events, and servicer receipts for the same research task.
Read full text and fault analysis → · Download Reliability Experiment v3 ↓
python3 cli.py memory-put
python3 cli.py submit
python3 cli.py run
python3 evaluate.py
python3 -m unittest discover -s . -p test_lab.py -vBy default, deterministic summary functions and synthetic documents are used; real operating trajectories and mechanical constraints are evaluated, and there is no semantic support to verify the real model.
Write a three-dimensional scoring rule and design an incorrect answer that can fool judges who only look at writing style.