Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Examine multi-tenant queuing, fair scheduling, backpressure and multi-dimensional quotas.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Idempotency, unknown outcomes, and task recovery →Control and completion criteria in agent loops →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
Reading implementation and trade-offs →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Fair scheduling, resource quotas, and backpressure across tenants
Preparatory concepts:Queuing and concurrency, Tenant resource isolation, Task recovery steps
The total concurrency limit protects system capacity, and tenant scheduling protects resource allocation. Both are indispensable. Fairness must be measured based on actual consumption and waiting. The number of requests or the global average delay cannot be regarded as an isolation guarantee.
Fast submission of 10,000 tasks can overload databases, models, and tools. Admission limits pending tasks, task size, and tenant budgets. Execution limits model concurrency, token rates, tool connections, and memory. Equal request counts can have very different costs.
It limits concurrent work, but large tenants may occupy every slot. Use tenant round-robin, weighted, or cost-aware scheduling, alongside tenant concurrency and global limits. Requeue long tasks at recoverable boundaries or isolate resource pools; remote writes cannot be arbitrarily preempted safely.
Tie priorities and weights to service objectives. Aging, minimum shares, or maximum-wait targets reduce starvation, but insufficient capacity still requires rejection or user choices. Amazon SQS fair queues uses tenant identity to mitigate noisy-neighbor waiting. It does not enforce per-tenant consumption rates or standard-queue ordering, so application quotas remain necessary.
Reduce admission when models throttle, tools fail, or indexes slow down. Use bounded backoff instead of unlimited retries. Separate bottleneck quotas keep one dependency from consuming every worker. Mixed-load tests measure small-tenant P95 wait, queue age, large-tenant throughput, actual cost, and starvation. A better global average may merely reflect more fast tasks from large tenants.
Use tenant admission quotas, fair scheduling, and global/downstream execution limits. Bound concurrency, tokens, calls, costs, and waiting as well as request counts. Long tasks yield at recoverable boundaries; handle cancelled and expired tasks promptly. Mixed-load tests measure small-tenant waits and large-tenant throughput rather than only global averages.
The quick return of the task ID by the submission interface does not mean that the backlog can be unlimited. Set the maximum number of pending tasks, task size and budget for the tenant, and return understandable retry or queuing information when the limit is exceeded. The execution layer limits concurrency based on the actual capacity of the model and tool. You cannot start 100,000 coroutines just because the queue can hold 100,000 items.
Tenant polling, weighted polling, or cost-based scheduling can be used, with weights determined by service levels. Long tasks are divided into recoverable steps and queued after the steps are completed to prevent one operation from occupying the Worker for a long time. Scheduling should avoid starvation and consider deadlines, but high priorities should not preempt other tenants indefinitely. If the estimated cost is inaccurate, it will be settled and adjusted based on actual usage.
When the model supplier is limited, the vector library is slow, or the tool gateway fails, the task fetching speed is reduced to avoid continuous failures and retry storms caused by queue consumers. Set independent quotas and circuit breakers for different resources. When canceling, clean up unstarted tasks and stop new actions. The issued write request enters the verification process as usual and cannot be considered completely canceled just by deleting it from the queue.
The stress test simultaneously adds a large tenant batch task, multiple small tenant interaction tasks, and downstream rate limiting. Observe P95 per tenant wait, queue age, completion rate, actual cost, and hunger. Gradually boost until you bottleneck, checking to see if the system is rejecting or slowing down rather than running out of memory. Global mean improvement cannot mask a tenant's complete unserviceability.
Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1Why is setting a global Semaphore not enough?
Drill down from protecting total capacity to tenant resource allocation.
It limits the total number of simultaneous runs and does not determine which tenant gets the quota. The 10,000 large tenants in the FIFO queue are still ahead of the small tenants, and long tasks may continue to fill the slots. Add tenant fair receipt, tenant concurrency and different resource quotas; the global quota of the multi-process system also needs to be shared and coordinated, and the intra-process Semaphore is not a full-service quota.
Follow this answer further
Level 2After adding Semaphore to each tenant, why do small tenants still have to wait for a long time?
After the parent increases the tenant limit, the queue collection sequence may still cause head congestion.
If the global FIFO first releases a large number of tasks that are blocked by tenant quotas, consumers may use these tasks to occupy slots or repeatedly idle before it is the turn of small tenants. First, runnable tenants are selected fairly, and then their tasks and resources are collected; if they fail to be collected, they are released in a timely manner. Queue scheduling and execution limits must be coordinated, and you cannot just stack two locks.
Follow this answer further
Level 3It is fair to poll based on the number of entries. For large tenants, each entry takes one hour, while for small tenants, it only takes ten seconds. Is this fair?
After fixing head-of-queue blocking, task cost differences require redefinition of fairness units.
Equal task counts do not imply fair resource use. Allocate by estimated execution cost or time slices, let long tasks yield at safe step boundaries, and adjust later quotas using actual cost. Use a separate pool or tighter concurrency for long steps that cannot be safely preempted. Evaluate small tenants' waiting times and resource shares, not just completed task counts.
Level 1What should I do if the Token cost cannot be accurately estimated in advance?
Fair request allocation also takes into account unknown and unequal costs.
Reserve according to conservative upper bound or estimated cost, limit the maximum output and number of tools in a single step, end settlement based on actual usage and adjust future estimates. If the reservation is exceeded, the quota must be added or stopped, and the unknown usage is left for reconciliation; scheduling can be fairly claimed according to the estimated cost first, and the actual difference will affect subsequent shares. Costs cannot be constrained based solely on the number of tasks.
Level 1Can a priority queue starve low-priority tasks?
Priority design handles starvation and capacity insatiability conditions.
Yes, persistent high-priority traffic may take away all resources forever. Set minimum shares, wait for entitlements or split pool capacity, limit high-priority admission, and define denial policies in the event of overload. You cannot simultaneously commit to unlimited high-priority throughput and bounded all low-priority waits; goals should be based on capacity and testing.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:Scheduling capacity remains unchanged, but actual service capacity decreases
Extended question:Can adding workers eliminate the backlog?
May only increase the concurrent failure and retry peak. Identify the model rate limiting signal, reduce the concurrency between receiving and the resource, avoid jitter, and isolate work that does not depend on the model. The entrance prompts to queue or reject tasks that exceed the budget, and continuously monitors the queue age; for capacity expansion, the bottleneck service must first be confirmed and the quota can be increased.
The principles that remain unchanged:Throughput is subject to the narrowest resource constraint, and backpressure needs to be based on actual downstream capacity rather than the number of Workers.
Changing conditions:There are also different latency targets within the same tenant
Extended question:Can per-tenant polling guarantee interactive waiting?
There is no guarantee that long tasks within a tenant will not block short tasks. Pools can be pooled by task type or prioritized within the tenant, and the lowest share reserved for the batch. Long tasks are given away in recoverable steps to limit the execution cost of one round; 10,000 high-priority subtasks cannot be split to bypass the original batch quota.
The principles that remain unchanged:Fair resource units should match task costs and service goals, and hierarchical scheduling still requires access constraints.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
Design scheduling and observation indicators for 1,000 long tasks for large tenants and 5 short tasks for small tenants.