Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Examine code base understanding, isolation execution, acceptance, regression and manual review.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Isolate resources and capabilities when executing code →Agent evaluation: outcomes, constraints, and evidence →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
Reading implementation and trade-offs →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Bind code changes to acceptance evidence
Preparatory concepts:Git differences, Test baseline, Build and dependencies
Evidence of code delivery must be bound to the actual commit and environment. Acceptance goals are determined by tasks and protected constraints, and the Agent cannot make a change appear complete by deleting tests, hiding failures, or reusing old results.
An empty-password login fix should reject empty passwords, preserve valid login and API compatibility, and avoid unrelated dependency changes. Read repository conventions, call chains, and tests; use an isolated workspace for a focused diff. Repository text cannot independently authorize deployment credentials.
If original tests fail, run relevant checks on both commits in the same environment. Distinguish pre-existing failures, environmental changes, and regressions. Do not dismiss failures or change unrelated code. Existing failures limit conclusions; the new task still needs targeted verification.
Changed files, lockfiles, or run parameters make old test results insufficient. Deliver commit details, commands, environment, actual results, and untested scope. PR submission and merging follow the authorized process. These are instructional assignments without new Java modifications or tests.
Start with a small verifiable task, repository conventions, call chains, and tests. Generate a focused diff in an isolated workspace. Models propose changes; runtime enforces permissions, budgets, tests, and review. Deliver changes, actual checks, unresolved risks, and reproduction steps. PR creation does not authorize merging; define approval boundaries for migrations, APIs, and dependencies.
Taking "fixing duplicate inventory deductions" as an example, we first require failure recurrence, target behavior, allowed modification range and compatibility conditions. Build a code index or search for entries, callers, transaction boundaries and tests on demand without blindly stuffing the entire repository into context. Record the baseline test status to avoid mistaking a failed test for a new regression or changing the test itself to no longer check for the original problem.
Each task has an independent branch or workspace, the execution environment does not have production credentials, and dependencies are downloaded and the network is controlled. Tools include reading, searching, editing, building, testing and difference checking. Permissions do not include deployment and merging by default. For large Maven or Gradle projects, you can run relevant module tests first, and then expand according to the scope of changes; the test selection needs to explain the reasons for coverage, and you cannot just select commands that are easy to pass.
Acceptance not only looks at compilation, but also regression testing, public API compatibility, SQL migration, exception paths and concurrent behavior. Save commands, exit codes, environment versions, and key output as evidence. Multiple rounds of repair are limited by budget; when recovering, the associated hash of code changes and run checks is retained. If the code changes, the old test passing status cannot be reused.
A PR describes the issue, final behavior, tests, and limitations, and provides minimal reproduction and rollback considerations. Reviewers should be able to see why the differences are valid. The technical leader will ask the Agent how to discover transaction defects, how to protect unmodifiable interfaces, and how to avoid deleting tests to "resolve failures", instead of just asking to demonstrate automatic code writing once.
Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1Is the Agent complete after deleting the failed test?
Task completion requires protection of acceptance criteria from being overwritten by the performer.
Doesn’t count. First check whether the deleted test protects the task goal or public behavior. It is forbidden to remove the assertion in exchange for a green light. If the requirements do change, you can update the test under a clear new contract and explain why the old assertion is no longer applicable, and still add effective coverage; the deletion result cannot be pretended to be a repair implementation.
Follow this answer further
Level 2The Agent says that the test is out of date and changes the expectation to new behavior. How to judge?
The parent question prohibits deletion of acceptance, and the child question adds retention testing but modifies the expected bypass.
Returning to tasks and public contracts, check whether the requirements authorize changing behavior and caller dependencies. Don't change expectations just because a candidate implementation produces new output; use independent acceptance samples or human review to confirm that the new behavior is correct. Test code is also a delivery difference and must be reviewed for purpose and coverage.
Follow this answer further
Level 3Testing and implementation are both generated by the same model. How to reduce common misunderstandings?
The parent needs to independently determine the new behavior and continue to pursue implementation and testing of common misunderstanding requirements.
Counterexamples are designed based on independent task constraints, existing public behavior tests are retained, and key assertions are checked by another round of review or manually; reference implementations or business fact checks are used when necessary. Multiple models by themselves do not guarantee independence. The key is that the sources of testing are different and cover wrong paths.
Level 1How to distinguish regression when existing tests fail?
Old systems often fail, and a fair before-and-after comparison is needed.
Lock the same environment and dependencies, rerun the baseline and candidate commits, and classify them by test and bug type. If the baseline also fails, it only means that the failure did not occur for the first time due to the candidate, and does not prove that the candidate does not have other regressions; unstable tests need to be diagnosed again or remain unknown, and the report can be verified in a range.
Level 1Can the previous test results be reused after modification?
Acceptance occurs at a certain version, and subsequent modifications may invalidate the evidence.
It cannot be reused directly. Which facts are supported by that check is only stated if the reported code, dependencies, configuration, and environment are still consistent. Run the affected checks after modification, and complex dependency changes need to expand the scope; indicate the items that have not been re-run, and do not mark the historical green light as currently passed.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:The difference extends from method fixes to persistent data structures.
Extended question:Are normal unit tests passing enough?
It is necessary to additionally verify the migration sequence, old data, compatible read-write and recovery solutions, and clarify the approval and actual operation scope. Independent test databases can simulate upgrades, and production migration cannot be claimed to be safe without verification in production; migration risks are delivered instead of automated execution.
The principles that remain unchanged:The evidence should cover the actual change contract, binding version and environment.
Changing conditions:Code review can be done, but dynamic results are missing.
Extended question:Can I submit a PR?
Reviewable differences, static checks and unrun items can be delivered within the scope of authorization, and reproduction commands and required environments can be provided. Do not make up the pass log; for high-risk modifications, first limit the delivery stage and explain the acceptance that needs to be made. The absence of execution evidence does not mean that all static judgments are worthless.
The principles that remain unchanged:Conclusions must not extend beyond a true examination.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
Design task contracts, tool permissions, failure reproduction and PR acceptance for Java duplicate deduction issues.