Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 37AdvancedSystem designAbout 18 minutes

Understand → Implement → Debug → Design

Fair scheduling, resource quotas, and backpressure across tenants

Examine multi-tenant queuing, fair scheduling, backpressure and multi-dimensional quotas.

Queuemulti-tenantback pressure

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Realize again

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

Reading implementation and trade-offs →

Will troubleshoot

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Able to choose

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Fair scheduling, resource quotas, and backpressure across tenants

Understand the core principles first

Preparatory concepts:Queuing and concurrency, Tenant resource isolation, Task recovery steps

The total concurrency limit protects system capacity, and tenant scheduling protects resource allocation. Both are indispensable. Fairness must be measured based on actual consumption and waiting. The number of requests or the global average delay cannot be regarded as an isolation guarantee.

Queue capacity differs from execution capacity

Fast submission of 10,000 tasks can overload databases, models, and tools. Admission limits pending tasks, task size, and tenant budgets. Execution limits model concurrency, token rates, tool connections, and memory. Equal request counts can have very different costs.

A global semaphore does not ensure fairness

It limits concurrent work, but large tenants may occupy every slot. Use tenant round-robin, weighted, or cost-aware scheduling, alongside tenant concurrency and global limits. Requeue long tasks at recoverable boundaries or isolate resource pools; remote writes cannot be arbitrarily preempted safely.

Fairness mechanisms have limits

Tie priorities and weights to service objectives. Aging, minimum shares, or maximum-wait targets reduce starvation, but insufficient capacity still requires rejection or user choices. Amazon SQS fair queues uses tenant identity to mitigate noisy-neighbor waiting. It does not enforce per-tenant consumption rates or standard-queue ordering, so application quotas remain necessary.

Propagate backpressure to admission

Reduce admission when models throttle, tools fail, or indexes slow down. Use bounded backoff instead of unlimited retries. Separate bottleneck quotas keep one dependency from consuming every worker. Mixed-load tests measure small-tenant P95 wait, queue age, large-tenant throughput, actual cost, and starvation. A better global average may merely reflect more fast tasks from large tenants.

Check understanding with a question

How can a large customer submit 10,000 Agent tasks without bringing down other tenants?

Use tenant admission quotas, fair scheduling, and global/downstream execution limits. Bound concurrency, tokens, calls, costs, and waiting as well as request counts. Long tasks yield at recoverable boundaries; handle cancelled and expired tasks promptly. Mixed-load tests measure small-tenant waits and large-tenant throughput rather than only global averages.

Realization and trade-offs

Entry restrictions are different from execution restrictions

The quick return of the task ID by the submission interface does not mean that the backlog can be unlimited. Set the maximum number of pending tasks, task size and budget for the tenant, and return understandable retry or queuing information when the limit is exceeded. The execution layer limits concurrency based on the actual capacity of the model and tool. You cannot start 100,000 coroutines just because the queue can hold 100,000 items.

Fair scheduling and long tasks

Tenant polling, weighted polling, or cost-based scheduling can be used, with weights determined by service levels. Long tasks are divided into recoverable steps and queued after the steps are completed to prevent one operation from occupying the Worker for a long time. Scheduling should avoid starvation and consider deadlines, but high priorities should not preempt other tenants indefinitely. If the estimated cost is inaccurate, it will be settled and adjusted based on actual usage.

Backpressure needs to penetrate the link

When the model supplier is limited, the vector library is slow, or the tool gateway fails, the task fetching speed is reduced to avoid continuous failures and retry storms caused by queue consumers. Set independent quotas and circuit breakers for different resources. When canceling, clean up unstarted tasks and stop new actions. The issued write request enters the verification process as usual and cannot be considered completely canceled just by deleting it from the queue.

How to prove that isolation is effective

The stress test simultaneously adds a large tenant batch task, multiple small tenant interaction tasks, and downstream rate limiting. Observe P95 per tenant wait, queue age, completion rate, actual cost, and hunger. Gradually boost until you bottleneck, checking to see if the system is rejecting or slowing down rather than running out of memory. Global mean improvement cannot mask a tenant's complete unserviceability.

Engineering deduction

scene
Interview hypothesis: Tenant A uploads 10,000 documents at a time, and tenant B only submits a question and answer once but waits for half an hour.
design decisions
Use tenant queues and authorized claims to independently limit batch and interactive task resources.
Verify target
B's waits meet the target, and A still has predictable throughput.
applicable boundary
The specific weight and capacity are determined by service level and stress testing.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Downstream model suddenly limits flow

Changing conditions:Scheduling capacity remains unchanged, but actual service capacity decreases

Extended question:Can adding workers eliminate the backlog?

Derivation and reference solutions

May only increase the concurrent failure and retry peak. Identify the model rate limiting signal, reduce the concurrency between receiving and the resource, avoid jitter, and isolate work that does not depend on the model. The entrance prompts to queue or reject tasks that exceed the budget, and continuously monitors the queue age; for capacity expansion, the bottleneck service must first be confirmed and the quota can be increased.

The principles that remain unchanged:Throughput is subject to the narrowest resource constraint, and backpressure needs to be based on actual downstream capacity rather than the number of Workers.

Mixing long batches with short interactions

Changing conditions:There are also different latency targets within the same tenant

Extended question:Can per-tenant polling guarantee interactive waiting?

Derivation and reference solutions

There is no guarantee that long tasks within a tenant will not block short tasks. Pools can be pooled by task type or prioritized within the tenant, and the lowest share reserved for the batch. Long tasks are given away in recoverable steps to limit the execution cost of one round; 10,000 high-priority subtasks cannot be split to bypass the original batch quota.

The principles that remain unchanged:Fair resource units should match task costs and service goals, and hierarchical scheduling still requires access constraints.

Easy to make mistakes

  • Unlimited queues and unlimited concurrency
  • Only according to the request limit
  • Only monitor the global average response

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Tenant quotas and concurrency limits can be set separately at the submission entrance and execution layer.
Intermediate and advanced signals
Can illustrate fair queuing, resource tiering, and backpressure.
Senior Signal
Design mixed-load stress tests and handle cost errors, starvation, and cancellations.

Hands-on verificationComplete on demand · Suggestions15 minutes

Design scheduling and observation indicators for 1,000 long tasks for large tenants and 5 short tasks for small tenants.

Expand acceptance requirements and checkpoints
  • Small tenants don’t have to wait indefinitely
  • Downstream current restriction reversely reduces consumption
  • Both cost and concurrency have bounds

Key inspections

  • Quota override submission and execution
  • Can provide fair scheduling and back pressure
  • Observe tail latency by tenant