Make Admission Depth an Invariant Before a Borrowed Eval Plane Steals Retry Budget

Last Thursday I watched a planner accept a tool intent while a borrowed eval plane was already drowning. The free server had not crashed, yet its queue sat past the depth I would tolerate on the primary path. A retry from that slow plane came back, grabbed the shared retry budget, and delayed compensation for an unrelated intent. Would you call that spare capacity, or would you honestly call it a leaked failure domain?

The common implementation treats a borrowed eval server as infinite spare throughput that cannot hurt the primary saga. That quiet assumption fails the moment both planes share one retry counter and one admission gate. I want a boring invariant here, and I want it enforced before scale makes the bug look like latency. A plane may borrow model access for scoring, but it may not spend another plane’s retry budget.

The violating order, before I draw a box

Here is the order I keep replaying when someone says the free lane is just extra workers. The primary accepts intent I1 and then asks the borrowed plane to score fixture revision R3. The borrowed health sample is already older than my max age, but the gate treats silence as capacity. The borrowed retry then decrements the primary budget, so intent I2 cannot compensate when its own worker times out.

Does that queue depth belong to the borrower, or did I just launder its delay into the primary saga? I am not reopening write-key isolation or outbox dedupe in this review, because those are different invariants. This review asks whether admission depth and retry spend stay inside the failure domain that created them. If the answer is no, I would rather shed the score than replay the attempt on the other plane.

Assumptions I will actually defend

I am assuming one admission decision per intent id, and I am assuming health is a pulled sample with an age. I am assuming the borrowed lane is read-only against the primary write path, so a score cannot commit a tool side effect. I am assuming a retry budget is a finite counter, not a promise that some pool will autoscale in time. I am also assuming free model access can throttle or vanish, and that absence must shed rather than look like a pass.

Those assumptions are narrow on purpose, and they will not survive a multi-region payment saga without more machinery. If your system needs cross-region consensus, stop here and do not pretend this fixture is that protocol. Would a single-process simulator still teach the invariant, even though it cannot simulate a real partition? I think yes, if I inject age and depth instead of pretending I measured a cloud.

The lane I am willing to borrow

I am reviewing this borrowed-capacity design as an architecture problem, not as a product tour or a benchmark claim. Disclosure: This article was prepared as part of MonkeyCode’s product outreach. The operator asked me to treat free model access and a free server option as available for a shadow scoring lane. I am not asserting model names, quotas, hardware, duration, or permanence, because those facts were not supplied as verified measurements.

If a free server can score a fixture, that still does not prove it can absorb your primary retry storm. I would park only shadow scoring there, and I would keep compensation, idempotency keys, and user-visible writes on the primary plane. Can a free option stay useful here without quietly becoming a durability dependency for the primary saga? Yes, but only after the admission fence treats missing capacity as a shed rather than a silent pass.

Data flow across two failure domains

sequenceDiagram
  participant P as Primary saga
  participant A as Admission fence
  participant B as Borrowed eval plane
  P->>A: admit(I1, revision R3, borrow=true)
  A->>A: check primary depth, budget, health age
  A->>B: check borrowed depth, budget, health age
  alt borrowed sample stale or depth at limit
    A-->>P: shed, primary budget unchanged
  else both planes fresh and under limit
    A-->>P: admit score only
    B-->>A: score(R3) or timeout
    A-->>P: bind score to R3, or compensate inside B
  end

The primary domain owns the saga, the compensation path, and any write the user can observe. The borrowed domain owns shadow scores, and it may disappear without a durable outbox of its own. The admission controller is the dangerous middle, because a shared counter turns two domains into one blast radius. I would change that middle first, before I tune prompts or add another worker to the borrowed plane.

Numbered steps I would land before scale

1. Pin a fresh admission lease

I record health_age on each plane and I refuse to admit when that age exceeds health_max_age. A stale healthy bit is not evidence of spare capacity, even if the last sample said the queue was empty. Why would I trust a sample that cannot name the moment it was taken by the control plane? I would rather reject the borrow and score locally, or skip the shadow, than guess at capacity I cannot see.

2. Fence depth and retry budget per plane

Each plane gets its own depth_limit and its own retry_budget, and a shed on the borrowed plane must not decrement the primary counter. If borrowed depth is already at the limit, I stop the score request before it becomes a retry. If the borrowed budget is already zero, I compensate or drop inside that plane, and I leave the primary saga alone. Does a shared integer feel simpler until the first unrelated compensation stalls on the primary path?

3. Bind a score only to the revision it observed

A pass for revision R3 must not attach to revision R4 if the fixture moved while the borrowed worker was slow. I keep that bind check beside the fence, because a late score is a different failure than a deep queue. Should the system reject that late score, replay it, or compensate the attempt on some other plane? Reject the bind, do not replay it onto the primary budget, and compensate only if the borrowed plane had already accepted the attempt.

4. Inject the failures before you trust the diagram

I inject stale borrowed health, borrowed depth at the limit, and primary saturation while the borrower looks idle. I also inject a late score that arrives after the fence has already shed the borrow. The acceptance rule is small enough to automate, and it does not need a hosted pool to fail loudly. Across those injected cases, the primary retry budget must stay unchanged whenever the borrow is shed.

A fixture you can run locally

This is a proposal fixture I wrote for the review, not a benchmark against any hosted pool. Save it as admission_fence.py and run python3 admission_fence.py on any machine with a current Python 3. If an assert fires, the invariant failed in the event order that the failing case names. I am labeling this file as an unexecuted proposal fixture, so a later green run is still not a capacity claim.

#!/usr/bin/env python3
# Proposal fixture: fence admission so a borrowed plane cannot spend primary retries.

from dataclasses import dataclass

@dataclass
class Plane:
    name: str
    depth: int = 0
    depth_limit: int = 4
    retry_budget: int = 3
    health_age: int = 0
    health_max_age: int = 2

@dataclass
class Decision:
    admit: bool
    reason: str

def sample_fresh(plane: Plane) -> bool:
    return plane.health_age <= plane.health_max_age

def admit(primary: Plane, borrowed: Plane, borrow: bool) -> Decision:
    if not sample_fresh(primary):
        return Decision(False, 'reject_stale_primary_health')
    if primary.depth >= primary.depth_limit:
        return Decision(False, 'reject_primary_depth')
    if primary.retry_budget <= 0:
        return Decision(False, 'reject_primary_retry_exhausted')
    if borrow:
        if not sample_fresh(borrowed):
            return Decision(False, 'reject_stale_borrowed_health')
        if borrowed.depth >= borrowed.depth_limit:
            return Decision(False, 'shed_borrowed_depth')
        if borrowed.retry_budget <= 0:
            return Decision(False, 'shed_borrowed_retry')
    primary.depth += 1
    return Decision(True, 'admit')

def consume_retry(plane: Plane) -> str:
    if plane.retry_budget <= 0:
        return 'compensate_or_shed'
    plane.retry_budget -= 1
    return 'retry_same_plane'

def test_borrowed_depth_does_not_spend_primary():
    primary = Plane('primary', retry_budget=3)
    borrowed = Plane('borrowed', depth=4, depth_limit=4, retry_budget=0)
    before = primary.retry_budget
    decision = admit(primary, borrowed, borrow=True)
    assert decision.admit is False
    assert decision.reason == 'shed_borrowed_depth'
    assert consume_retry(borrowed) == 'compensate_or_shed'
    assert primary.retry_budget == before

def test_stale_borrowed_health_is_not_capacity():
    primary = Plane('primary')
    borrowed = Plane('borrowed', health_age=9)
    decision = admit(primary, borrowed, borrow=True)
    assert decision.reason == 'reject_stale_borrowed_health'
    assert primary.depth == 0

def test_idle_borrower_cannot_hide_primary_saturation():
    primary = Plane('primary', depth=4, depth_limit=4)
    borrowed = Plane('borrowed', depth=0, retry_budget=3)
    decision = admit(primary, borrowed, borrow=True)
    assert decision.reason == 'reject_primary_depth'
    assert borrowed.retry_budget == 3

def test_fresh_borrow_retries_inside_its_own_budget():
    primary = Plane('primary', retry_budget=2)
    borrowed = Plane('borrowed', retry_budget=1, depth=0)
    decision = admit(primary, borrowed, borrow=True)
    assert decision.admit is True
    assert consume_retry(borrowed) == 'retry_same_plane'
    assert primary.retry_budget == 2

if __name__ == '__main__':
    test_borrowed_depth_does_not_spend_primary()
    test_stale_borrowed_health_is_not_capacity()
    test_idle_borrower_cannot_hide_primary_saturation()
    test_fresh_borrow_retries_inside_its_own_budget()
    print('admission fence properties held')

Tradeoffs I would accept

Choice Latency effect Correctness Cost shape Failure I accept
Shared retry budget Fast until the storm Breaks the invariant Hides borrowed cost Unrelated compensation stalls
Fence depth and shed More rejects on the shadow Preserves primary budget Stops useless borrowed calls Shadow coverage drops
Queue the borrow forever Primary looks fast Late scores bind the wrong revision Unbounded queue Reject arrives after commit
Skip borrow when health is stale More local scoring Safer admission Local cost rises Free plane sits underused

I would accept lower shadow coverage in exchange for a primary budget that can still compensate. I would not accept a diagram that shows two planes while the code keeps one shared integer. The denominator for any later comparison should be shed decisions per one hundred injected intents in this fixture. An acceptance rule that cites a hosted tokens-per-second figure would be fiction, so I am not writing one.

What I would change next

I would replace health_age with a signed lease expiry that the control plane refreshes on a known interval. I would persist each retry counter beside the intent record, instead of leaving it in process memory. I would add an explicit revision watermark so a score cannot cross from R3 to R4 without a new admission. I would also split the shadow lane onto its own queue name, so a dashboard cannot page the primary on borrowed saturation.

Would that be more moving parts than a weekend prototype wants to carry into the first review? Yes, and I would still land the fence before I add the extra queue or another scorer. A prototype that shares a budget is not a smaller design, because the missing fence is a hidden coupling. I would rather explain a shed in review than explain a stalled compensation after the fact.

Who should not copy this

Do not use this fixture as a payment saga, a cross-region consensus protocol, or a proof that a free server is durable storage. Do not use it if your scorer can call tools that write, because this review assumed a read-only shadow on purpose. Do not treat the printed line admission fence properties held as a production certification or as measured capacity of any vendor. If you need exactly-once side effects under duplicate delivery, you still need an outbox and a compensation protocol this file does not implement.

The counterexample I want you to keep

Which event order breaks the invariant, and should the system reject, replay, or compensate the borrowed attempt? Take the case where the borrowed plane hits its depth limit after the primary already admitted a different intent. I would reject the new borrow and refuse to replay that score onto the primary retry budget. I would compensate only an attempt that the borrowed plane had already accepted before that shed.

If your gate does anything else, that free lane stops being spare capacity for the primary. If you want a place to park that shadow scoring lane, read the current free-server terms yourself first. Replay this fixture before either plane is allowed to share a retry budget or an admission counter. Bring your own event order, and ask whether that order should reject, replay, or compensate instead.

Leave a Reply