Skip to content

[None][fix] Declare attention runtime-workspace bytes/token as a backend contract - #16432

Open
eopXD wants to merge 1 commit into
NVIDIA:mainfrom
eopXD:attention-workspace-reservation-contract
Open

[None][fix] Declare attention runtime-workspace bytes/token as a backend contract#16432
eopXD wants to merge 1 commit into
NVIDIA:mainfrom
eopXD:attention-workspace-reservation-contract

Conversation

@eopXD

@eopXD eopXD commented Jul 15, 2026

Copy link
Copy Markdown
Collaborator

Description

Follow-up to #16399 (fp8 context-MLA attention-workspace reservation, nvbugs/6368562), now rebased
onto main with #16399 merged.

Motivation. #16399 reserves KV-cache headroom for the fp8 context-MLA attention workspace and caps
summed attended KV length in the scheduler. The accounting works, but it is keyed on a model-config
check (is this MLA + fp8 KV?) rather than on the backend that actually allocates the buffer. Two
reviewers flagged this on #16399:

  • @SimengLiu-nv: "does this function only apply to TRTLLM MLA or other sources as well like flashinfer?
    If only to TRTLLM MLA, it would be necessary to check the attention backend."
  • @QiJune: "Please apply the reserve and admission cap only when block reuse is enabled, chunked prefill
    is disabled, and the selected backend can reach the dense TRTLLM full-gather path."

#16399 landed the first two conditions of QiJune's list; this PR lands the third. It was deferred by
agreement, not dropped.

The underlying failure mode is structural, not MLA-specific: the KV-cache estimator profiles peak memory
against an empty cache and hands the rest to the KV pool, so any attention backend that stages a
workspace sized by a runtime quantity the profiling forward does not drive to its serving maximum (here
total_kv_len, decoupled from max_num_tokens by KV-cache reuse) is under-reserved and can OOM. A
future backend would have to rediscover and re-thread the same estimator + cost-rate pieces, and
silently re-introduce the same OOM if it missed one.

Change. Lift the accounting into a declared contract on the attention backend, so future backends
inherit the accounting instead of the OOM:

  • AttentionBackend.runtime_workspace_bytes_per_token(model_config, mapping) -> int (default 0) — a
    backend declares the per-token cost of any workspace it stages whose size scales with such a runtime
    quantity.
  • TrtllmAttention declares the fp8 context-MLA workspace, still sized by the single C++ source of
    truth (AttentionOp::contextMlaWorkspaceBytesPerToken, via nanobind), and keeping [https://nvbugs/6368562][fix] Reserve fp8 context-MLA attention workspace in KV cache estimation #16399's
    runtime-matched sparse gate verbatim (dsa/deepseek_v4 on SM 100/103 with the short-seq MHA fallback
    off — not "a sparse config exists").
  • The estimator resolves the declaration through the model's selected backend via
    get_attention_workspace_bytes_per_token(). The reserve/cap math is unchanged.
  • Documented as a contract in ATTENTION_DEVELOPER_GUIDE.md (required reading) — §2.3, §3.2.3, §4.2.

This is not a pure refactor. A model that resolves to a non-TRTLLM backend now correctly reserves
nothing, where main charges it the MLA rate off a model-config check and shrinks the KV pool for a
buffer that backend never allocates. That is the behavior change the two reviews asked for. The active
fp8-MLA-on-TRTLLM path — the nvbugs/6368562 repro — is unchanged.

Scope of the contract. Deliberately a scalar per-token rate, not a typed driver/reservation
abstraction. Only the rate is generalized; the driving quantity (total_kv_len), the reservation gate
(get_mla_context_workspace_kv_len_cap) and the cap plumbing (kv_cache_manager.fp8_ctx_mla_kv_len_cap,
KvCacheConfig.fp8_context_mla_kv_len_cap) remain MLA-named, because there is one driver today and the
scheduler's cap is specific to it. A backend with a different driving quantity introduces it then,
alongside the enforcement it needs — a richer type now would be unused scaffolding.

Note that after #16399's review redesign (carrying the admission cap from the estimator onto the KV
manager instead of re-deriving it from pool layout), the scheduler no longer reads the per-token rate at
all. The contract therefore has exactly one consumer: the estimator. py_executor.py's only delta here
is the two review follow-ups below.

Additionally: two follow-ups from #16399 review threads

Both threads were marked resolved on #16399 without a code change landing. Both are in
PyExecutor._get_ctx_mla_kv_len_cap, the cap reader this contract feeds, so they are folded in here
rather than left dangling:

  • A carried cap of exactly 0 was collapsed to None ("no cap") by a truthiness check
    (int(carried) if carried else None), inverting admission control for precisely the tightest-budget
    case the reservation exists to protect. Now compares against None. Flagged by CodeRabbit; new test
    test_ctx_cap_zero_is_a_cap_not_no_cap asserts both the read and that the trim still enforces it
    (keeping the first request as the forward-progress guard).
  • getattr(self, "is_warmup", False)self.is_warmup — it is a real property on PyExecutor
    (py_executor.py:1212), so the defensive getattr is unnecessary. Flagged (non-blocking) by
    @pengbowang-nv. The remaining getattr on kv_cache_manager is kept deliberately and now carries a
    comment saying why: managers not built by the estimator never carry the attribute.

Known follow-ups (not in this PR)

  • The BF16 full_k / full_kv full-gather buffers on the reuse path are still unaccounted for
    (@QiJune's [https://nvbugs/6368562][fix] Reserve fp8 context-MLA attention workspace in KV cache estimation #16399 thread, deferred by agreement), along with the high-fanout shared-prefix memory test
    he asked to be tracked.
  • try_prepare_estimation disables estimation for context parallelism and the VANILLA backend without
    setting _skip_est, so configure_kv_cache_capacity never runs and no cap is installed — fail-open
    rather than fail-closed. Narrow (neither path reaches fp8 context-MLA today) but worth closing.

Test Coverage

  • tests/unittest/_torch/executor/test_mla_workspace_reserve.py — retargeted to the new resolver. The
    non-MLA test now exercises the full resolve-backend path (get_attention_backend → backend classmethod
    0); a new test_workspace_bytes_zero_for_backend_without_declaration covers a backend that
    inherits the default 0 for a model the TRTLLM backend would charge for (uses VANILLA, which
    always resolves — FLASHINFER silently falls back to TRTLLM when flashinfer is absent). Plus
    test_ctx_cap_zero_is_a_cap_not_no_cap for the carried-zero fix. All of [https://nvbugs/6368562][fix] Reserve fp8 context-MLA attention workspace in KV cache estimation #16399's existing coverage is
    preserved.
  • tests/unittest/_torch/executor/test_kv_cache_estimation.py — patch target renamed to the resolver.

PR Checklist

  • PR description clearly explains what and why.
  • PR follows TRT-LLM coding guidelines; pre-commit run locally against the PR diff range (green).
  • Test cases provided for the new code paths.
  • No public API changes (the backend method is internal).
  • No new dependencies.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

@coderabbitai

coderabbitai Bot commented Jul 15, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Adds a shared FP8 context-MLA workspace estimator, exposes it to Python backends, reserves corresponding KV-cache headroom, and caps scheduled context requests by attended KV length. Tests cover reuse accounting, trimming behavior, disabled caps, and non-MLA configurations.

Changes

FP8 context-MLA workspace accounting

Layer / File(s) Summary
Workspace estimator and binding
cpp/tensorrt_llm/common/attentionOp.*, cpp/tensorrt_llm/nanobind/thop/bindings.cpp
Adds the shared per-token K/V staging-byte estimator, synchronizes runtime buffer sizing with it, and exposes it through nanobind.
Backend workspace contract
tensorrt_llm/_torch/attention_backend/interface.py, tensorrt_llm/_torch/attention_backend/trtllm.py, tensorrt_llm/_torch/modules/ATTENTION_DEVELOPER_GUIDE.md
Adds the backend workspace hook, implements FP8 context-MLA detection and sizing, and documents the memory-accounting contract.
KV-cache capacity reservation
tensorrt_llm/_torch/pyexecutor/_util.py
Queries backend workspace requirements and adjusts KV-cache capacity to reserve runtime workspace headroom.
Context scheduling admission cap
tensorrt_llm/_torch/pyexecutor/py_executor.py, tests/unittest/_torch/executor/test_mla_workspace_reserve.py
Computes reuse-aware attended KV length, trims context requests to the configured cap, and tests the cap and workspace-resolution behavior.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Sequence Diagram(s)

sequenceDiagram
  participant AttentionBackend
  participant AttentionOp
  participant KvCacheCreator
  participant PyExecutor

  AttentionBackend->>AttentionOp: obtain FP8 context-MLA bytes per token
  AttentionOp-->>AttentionBackend: return K/V staging cost
  KvCacheCreator->>AttentionBackend: resolve workspace reservation
  AttentionBackend-->>KvCacheCreator: return bytes per token
  KvCacheCreator->>PyExecutor: establish KV capacity and attended-KV cap
  PyExecutor-->>PyExecutor: trim context requests exceeding the cap
Loading

Possibly related PRs

Suggested reviewers: sunnyqgg, yihwang-nv, cascade812, pengbowang-nv

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 36.67% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes the main change: declaring attention runtime-workspace bytes per token as a backend contract.
Description check ✅ Passed The description explains the motivation, implementation, behavior change, tests, scope, known follow-ups, and checklist items in the required sections.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (3)
cpp/tensorrt_llm/common/attentionOp.h (1)

61-67: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Document the new public API with Doxygen.

AttentionOp::contextMlaWorkspaceBytesPerToken is a new public interface, but its declaration uses ordinary // comments and does not document its parameters or return value. Use a Doxygen comment here.

As per coding guidelines, use Doxygen comments for new interfaces.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@cpp/tensorrt_llm/common/attentionOp.h` around lines 61 - 67, Replace the
ordinary comment immediately preceding
AttentionOp::contextMlaWorkspaceBytesPerToken with a Doxygen comment that
documents the method’s purpose, every parameter, and its returned byte count.
Preserve the existing sizing behavior and shared-source-of-truth description.

Source: Coding guidelines

tests/unittest/_torch/executor/test_mla_workspace_reserve.py (1)

67-71: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Coverage gap: cap-derivation and KV-budget-split logic are untested.

Coverage of _context_attended_kv_len and _cap_context_by_total_kv_len is solid, but two related pieces of new behavior have no test in this file:

  • tensorrt_llm/_torch/pyexecutor/py_executor.py::PyExecutor._get_ctx_mla_kv_len_cap — the _make_executor helper here bypasses it entirely by pre-setting _ctx_mla_kv_len_cap, so the actual blocks_in_primary_pool * tokens_per_block computation and the w > 0 gating are never exercised.
  • tensorrt_llm/_torch/pyexecutor/_util.py::KvCacheCreator.configure_kv_cache_capacity — the new cap = budget / (k + w) reservation-split branch has no unit coverage.

Both would need mocking (kv_cache_manager attributes / get_attention_workspace_bytes_per_token) similar to the pattern already used for _make_executor, so this is a reasonable, low-effort follow-up rather than a blocker — flagging for completeness. As per path instructions, calling out coverage gaps with concrete file names for QA follow-up.

Also applies to: 106-112

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/unittest/_torch/executor/test_mla_workspace_reserve.py` around lines 67
- 71, Add focused tests for the uncovered cap-derivation and KV-budget-split
branches. Extend the `_make_executor`-related tests to exercise
`PyExecutor._get_ctx_mla_kv_len_cap`, including `blocks_in_primary_pool *
tokens_per_block` and the `w > 0` gating, using mocked `kv_cache_manager`
attributes; add `KvCacheCreator.configure_kv_cache_capacity` coverage for the
`cap = budget / (k + w)` reservation split with a mocked
`get_attention_workspace_bytes_per_token`.

Source: Path instructions

tensorrt_llm/_torch/attention_backend/trtllm.py (1)

1296-1299: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win

Reuse the shared MLA predicate here.
tensorrt_llm._torch.pyexecutor.config_utils.is_mla() already checks both kv_lora_rank and qk_rope_head_dim; this guard only checks kv_lora_rank, so a malformed config can drift past the check and hit the later config.qk_rope_head_dim access. Importing the shared helper here is safe.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tensorrt_llm/_torch/attention_backend/trtllm.py` around lines 1296 - 1299,
Update the MLA guard in the relevant attention backend method to use the shared
config_utils.is_mla() predicate instead of checking kv_lora_rank directly, and
import that helper. Preserve the existing early return of 0 for non-MLA
configurations while ensuring both required MLA fields are validated before
later qk_rope_head_dim access.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tensorrt_llm/_torch/pyexecutor/py_executor.py`:
- Around line 5108-5133: Reset the cached _ctx_mla_kv_len_cap in
_maybe_rebalance_kv_pools immediately after a successful mgr.impl.adjust() so
the next _get_ctx_mla_kv_len_cap() recomputes blocks_in_primary_pool *
tokens_per_block using the rebalanced KV pool. Do not invalidate the cache when
adjust fails or is not performed.

---

Nitpick comments:
In `@cpp/tensorrt_llm/common/attentionOp.h`:
- Around line 61-67: Replace the ordinary comment immediately preceding
AttentionOp::contextMlaWorkspaceBytesPerToken with a Doxygen comment that
documents the method’s purpose, every parameter, and its returned byte count.
Preserve the existing sizing behavior and shared-source-of-truth description.

In `@tensorrt_llm/_torch/attention_backend/trtllm.py`:
- Around line 1296-1299: Update the MLA guard in the relevant attention backend
method to use the shared config_utils.is_mla() predicate instead of checking
kv_lora_rank directly, and import that helper. Preserve the existing early
return of 0 for non-MLA configurations while ensuring both required MLA fields
are validated before later qk_rope_head_dim access.

In `@tests/unittest/_torch/executor/test_mla_workspace_reserve.py`:
- Around line 67-71: Add focused tests for the uncovered cap-derivation and
KV-budget-split branches. Extend the `_make_executor`-related tests to exercise
`PyExecutor._get_ctx_mla_kv_len_cap`, including `blocks_in_primary_pool *
tokens_per_block` and the `w > 0` gating, using mocked `kv_cache_manager`
attributes; add `KvCacheCreator.configure_kv_cache_capacity` coverage for the
`cap = budget / (k + w)` reservation split with a mocked
`get_attention_workspace_bytes_per_token`.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 4157f3a6-9cad-4fa5-892c-67e8a22b83bc

📥 Commits

Reviewing files that changed from the base of the PR and between 0846183 and e18a63b.

📒 Files selected for processing (10)
  • cpp/tensorrt_llm/common/attentionOp.cpp
  • cpp/tensorrt_llm/common/attentionOp.h
  • cpp/tensorrt_llm/nanobind/thop/bindings.cpp
  • tensorrt_llm/_torch/attention_backend/interface.py
  • tensorrt_llm/_torch/attention_backend/trtllm.py
  • tensorrt_llm/_torch/modules/ATTENTION_DEVELOPER_GUIDE.md
  • tensorrt_llm/_torch/pyexecutor/_util.py
  • tensorrt_llm/_torch/pyexecutor/py_executor.py
  • tests/integration/test_lists/waives.txt
  • tests/unittest/_torch/executor/test_mla_workspace_reserve.py
💤 Files with no reviewable changes (1)
  • tests/integration/test_lists/waives.txt

Comment on lines +5108 to +5133
def _get_ctx_mla_kv_len_cap(self):
"""Cap on the summed context attended-KV length (total_kv_len) per forward step.

Equals the KV pool's primary-pool token capacity, which is exactly what the estimator reserved the
fp8 context-MLA attention workspace for (`max_tokens = budget / (k + w)`). Returns None (no cap) for
non-fp8-MLA models — those reserved no workspace, so this per-forward constraint must not alter their
scheduling. Computed once and cached.
"""
cap = getattr(self, "_ctx_mla_kv_len_cap", "unset")
if cap != "unset":
return cap
# Lazy import: _util imports py_executor at module scope, so a top-level import here is circular.
from ._util import get_attention_workspace_bytes_per_token
cap = None
w = get_attention_workspace_bytes_per_token(
self.model_engine.model.model_config, self.dist.mapping)
if w > 0:
blocks = getattr(self.kv_cache_manager, "blocks_in_primary_pool",
None)
tokens_per_block = getattr(self.kv_cache_manager,
"tokens_per_block", None)
if blocks and tokens_per_block:
cap = int(blocks) * int(tokens_per_block)
self._ctx_mla_kv_len_cap = cap
return cap

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Locate the relevant methods and nearby context.
python3 - <<'PY'
from pathlib import Path
path = Path("tensorrt_llm/_torch/pyexecutor/py_executor.py")
lines = path.read_text().splitlines()
targets = ["def _get_ctx_mla_kv_len_cap", "def _maybe_rebalance_kv_pools", "def _context_attended_kv_len", "def _cap_context_by_total_kv_len", "def _schedule"]
for t in targets:
    for i, line in enumerate(lines, 1):
        if t in line:
            start = max(1, i - 20)
            end = min(len(lines), i + 80)
            print(f"\n=== {t} @ line {i} ===")
            for j in range(start, end + 1):
                print(f"{j:5d}: {lines[j-1]}")
            break
PY

Repository: NVIDIA/TensorRT-LLM

Length of output: 30267


Invalidate the cached context-MLA KV cap after KV pool rebalance. _get_ctx_mla_kv_len_cap() caches blocks_in_primary_pool * tokens_per_block once, but _maybe_rebalance_kv_pools() can change the primary pool via mgr.impl.adjust(). Reset _ctx_mla_kv_len_cap after a successful adjust so scheduling tracks the current capacity.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tensorrt_llm/_torch/pyexecutor/py_executor.py` around lines 5108 - 5133,
Reset the cached _ctx_mla_kv_len_cap in _maybe_rebalance_kv_pools immediately
after a successful mgr.impl.adjust() so the next _get_ctx_mla_kv_len_cap()
recomputes blocks_in_primary_pool * tokens_per_block using the rebalanced KV
pool. Do not invalidate the cache when adjust fails or is not performed.

…end contract

The fp8 context-MLA workspace reservation (nvbugs/6368562, NVIDIA#16399) was threaded
imperatively through the KV-cache estimator, keyed on a model-config check
specific to MLA rather than on the backend that allocates the buffer. Two
reviewers flagged this on NVIDIA#16399: the reserve fires off a model-level MLA check,
so a model running a backend that never stages the buffer is still charged for
it -- shrinking the KV pool for a workspace it will not allocate.

Lift the accounting into a declared contract on the attention backend:

- AttentionBackend.runtime_workspace_bytes_per_token(model_config, mapping)
  returns the per-token bytes to reserve for a workspace the backend stages whose
  size scales with a runtime quantity the profiling forward does not drive to its
  serving maximum. Default 0 -- correct for every backend but fp8 context-MLA.
- TrtllmAttention declares the fp8 context-MLA K/V dequant workspace, still sized
  by the single C++ source of truth (contextMlaWorkspaceBytesPerToken) and
  keeping NVIDIA#16399's runtime-matched sparse gate (dsa/deepseek_v4 on SM 100/103
  with the short-seq MHA fallback off).
- The estimator resolves the declaration through the model's selected backend via
  get_attention_workspace_bytes_per_token(). The reserve/cap math is unchanged;
  what changes is that a non-TRTLLM backend now correctly reserves nothing.
- Document the contract in ATTENTION_DEVELOPER_GUIDE.md (required reading) so a
  new backend inherits the accounting instead of the OOM.

The contract is deliberately a scalar per-token rate, not a typed
driver/reservation abstraction: there is one driving quantity today
(total_kv_len) and the scheduler's cap is specific to it, so a richer type would
be unused scaffolding. A backend with a different driver introduces it then,
alongside the enforcement it needs.

Also carries two follow-through fixes from NVIDIA#16399 review threads that were
resolved without a code change, both in the cap reader this contract feeds:

- A carried cap of exactly 0 was collapsed to None ("no cap") by a truthiness
  check, inverting admission control for the tightest-budget case it exists to
  protect. Compare against None instead.
- is_warmup is a real property on PyExecutor, so the defensive getattr is
  unnecessary.

Signed-off-by: Yueh-Ting Chen <yuehtingc@nvidia.com>
@eopXD
eopXD force-pushed the attention-workspace-reservation-contract branch from e18a63b to 5b0ff6b Compare August 3, 2026 06:06
@eopXD eopXD changed the title [None][chore] Declare attention runtime-workspace bytes/token as a backend contract [None][fix] Declare attention runtime-workspace bytes/token as a backend contract Aug 3, 2026
@eopXD
eopXD marked this pull request as ready for review August 3, 2026 06:11
@eopXD
eopXD requested a review from a team as a code owner August 3, 2026 06:11
@eopXD
eopXD requested a review from VALLIS-NERIA August 3, 2026 06:11
@eopXD
eopXD requested a review from lowsfer August 3, 2026 06:11
@eopXD

eopXD commented Aug 3, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63372 [ run ] triggered by Bot. Commit: 5b0ff6b Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63372 [ run ] completed with state FAILURE. Commit: 5b0ff6b
/LLM/main/L0_MergeRequest_PR pipeline #51356 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@BowenFu

BowenFu commented Aug 3, 2026

Copy link
Copy Markdown

Reviewed the full change. The refactor is sound and the moved TRTLLM body is behavior-preserving; two things worth calling out.

The backend-resolution delta is a silent fix that the description doesn't mention. Previously the fp8 MLA cost was computed regardless of attn_backend; now non-TRTLLM backends inherit the default 0. FlashInferAttention and VanillaAttention both report support_mla() == True, so an MLA + fp8-KV model on either used to get a reserve carved out of the KV pool — but neither touches the AttentionOp dequant workspace (vanilla goes through SDPA, flashinfer through its own paged-MLA path). So those configs were being over-reserved and now recover that pool. That's the right outcome, but it's a real behavior change for existing FLASHINFER/VANILLA MLA deployments and is worth a line in the PR body. get_attention_backend falling back to TrtllmAttention for unavailable/unknown names is what keeps the declaration matched to the backend create_attention actually builds — good.

cap == 0 is now enforced, with no operator-visible signal. Changing if carried to if carried is not None is right — collapsing 0 into "no cap" disabled admission control exactly where it's most needed. But a carried 0 means _cap_context_by_total_kv_len defers everything past the forward-progress request on every iteration, and the only trace is logger.debug. A deployment whose budget lands there sees context throughput collapse to one request per step with nothing in the log at default level. Consider a warn-once when the cap is first applied (or when it's 0), rather than per-iteration debug.

Not blocking on either. Holding my approval only on the red L0 on 5b0ff6b and on being first approver here — this touches the KV-cache estimator, the attention backend contract and the scheduler admission path, so I'd rather not be the only sign-off. Happy to approve once CI is green and someone closer to the estimator has looked.

@eopXD

eopXD commented Aug 3, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63447 [ run ] triggered by Bot. Commit: 5b0ff6b Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63447 [ run ] completed with state SUCCESS. Commit: 5b0ff6b
/LLM/main/L0_MergeRequest_PR pipeline #51418 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@eopXD

eopXD commented Aug 4, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63602 [ run ] triggered by Bot. Commit: 5b0ff6b Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63602 [ run ] completed with state FAILURE. Commit: 5b0ff6b
/LLM/main/L0_MergeRequest_PR pipeline #51564 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants