[NVBUG-6448152][test] Improve consensus efficiency in Python transceiver (DO NOT REVIEW YET) - #16766
Conversation
|
/bot run --disable-fail-fast --stage-list "DGX_H100-PyTorch-1, DGX_H100-PyTorch-2, DGX_H100-PyTorch-3, DGX_H100-PyTorch-4, DGX_H100-PyTorch-5, DGX_H100-PyTorch-6" |
|
PR_Github #61137 [ run ] triggered by Bot. Commit: |
|
PR_Github #61137 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "DGX_H100-PyTorch-1, DGX_H100-PyTorch-2, DGX_H100-PyTorch-3, DGX_H100-PyTorch-4, DGX_H100-PyTorch-5, DGX_H100-PyTorch-6" |
|
PR_Github #61190 [ run ] triggered by Bot. Commit: |
|
/bot run --disable-fail-fast --stage-list "DGX_H100-PyTorch-1, DGX_H100-PyTorch-2, DGX_H100-PyTorch-3, DGX_H100-PyTorch-4, DGX_H100-PyTorch-5, DGX_H100-PyTorch-6" |
9f71926 to
da52d66
Compare
|
/bot run --disable-fail-fast --stage-list "DGX_H100-PyTorch-1, DGX_H100-PyTorch-2, DGX_H100-PyTorch-3, DGX_H100-PyTorch-4, DGX_H100-PyTorch-5, DGX_H100-PyTorch-6" |
|
PR_Github #61363 [ run ] triggered by Bot. Commit: |
|
PR_Github/16766-9f71926 #61190 was force-killed by a newer pipeline run. |
|
PR_Github #61363 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "DGX_H100-PyTorch-1, DGX_H100-PyTorch-2, DGX_H100-PyTorch-3, DGX_H100-PyTorch-4, DGX_H100-PyTorch-5, DGX_H100-PyTorch-6" |
|
PR_Github #61381 [ run ] triggered by Bot. Commit: |
|
PR_Github #61381 [ run ] completed with state
|
fceed0f to
1301fb8
Compare
|
/bot run --disable-fail-fast --stage-list "DGX_H100-PyTorch-2, DGX_H100-PyTorch-3, DGX_H100-PyTorch-4, DGX_H100-PyTorch-5" |
|
PR_Github #61628 [ run ] triggered by Bot. Commit: |
|
PR_Github #61628 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "DGX_H100-PyTorch-1, DGX_H100-PyTorch-4" |
|
PR_Github #61655 [ run ] triggered by Bot. Commit: |
|
PR_Github #61655 [ run ] completed with state |
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #61671 [ run ] triggered by Bot. Commit: |
|
PR_Github #61671 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #61674 [ run ] triggered by Bot. Commit: |
|
PR_Github #61674 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #62347 [ run ] triggered by Bot. Commit: |
|
PR_Github #62347 [ run ] completed with state
|
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #62609 [ run ] triggered by Bot. Commit: |
|
PR_Github #62609 [ run ] completed with state |
Terminal generation-first Python throughput evidenceThe exact targeted run for head Frozen configuration:
Results:
Interpretation and remaining caveatsThis proves that the integrated generation-first Python-transceiver candidate completes the workload end to end. The result is 8.74 tok/s (1.89%) below the historical 463.66 tok/s context-first result, so it demonstrates restored generation-first functionality at approximately the historical Python throughput magnitude; it is not evidence of a raw throughput improvement over that unmatched context-first run. The stronger per-request terminal-agreement criterion was not fully observed before harness teardown. READY completed 256/256 on every context rank, but the shutdown summaries contained 251/256 terminal commits. Four tail rounds lacked rank three's final terminal vote and one lacked all final votes. All client requests had already completed successfully, and consensus/service shutdown was clean, so this does not invalidate the official client-throughput measurement. It does mean this run must not be described as proving 256 terminal commits. A pre-teardown drain/assertion is needed for that stronger evidence. The validated head contains two behavioral corrections after the production-equivalent boundary that are not yet present in the official Python production draft:
Those changes require productionization and review before this evidence can be attributed to the official production draft alone. |
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com> # Conflicts: # tests/integration/defs/perf/test_perf_sanity.py
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #62847 [ run ] triggered by Bot. Commit: |
|
PR_Github #62847 [ run ] completed with state
|
Paired terminal-consensus A/B — censored runThe exact paired A/B trigger for head Observed evidence:
Interpretation: this run is censored by a real transfer-timeout/lifecycle failure plus a fail-fast gap. It says nothing about the performance effect of asynchronous versus blocking terminal PP agreement. The unchanged head will not be rerun. The next attempt must first correct the experiment's transfer-timeout contract for a roughly 77-minute, concurrency-256 generation-first arm and retain strict fail-closed evidence requirements, then pass review before one new exact-stage trigger. |
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-2" |
|
PR_Github #63019 [ run ] triggered by Bot. Commit: |
|
PR_Github #63019 [ run ] completed with state
|
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Split paired terminal-consensus A/B — invalid runThe exact split paired-A/B trigger for head Both raw workloads completed 256/256 requests with zero client failures:
Those numbers must not be compared. Exact CTX evidence shows that both stages actually ran The run also exposed three evidence-harness defects after client completion:
There is therefore no accepted throughput delta, no valid asynchronous arm, and no The failed head is preserved on Replacement head |
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-2" |
1 similar comment
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-2" |
|
PR_Github #63066 [ run ] triggered by Bot. Commit: |
|
PR_Github #63066 [ run ] completed with state
|
Repaired terminal-consensus throughput A/B — censoredThe exact repaired trigger is terminal. Both arms used their intended CTX modes and completed 256/256 client requests with zero HTTP failures:
That raw difference is not interpretable. Every async CTX rank recorded only 255/256 terminal votes and commits, so the strict evidence check failed. The activation count was 256 in both arms, but the prompt-content digests differed across the two allocations, and neither downloaded result archive contained the required Therefore this run provides no valid terminal-rendezvous throughput conclusion. It will not be rerun unchanged. The diagnostic branch is preserved and this TEST ONLY draft is being closed unmerged. See the terminal CI report. |
Purpose
TEST ONLY — DO NOT REVIEW OR MERGE.
This draft runs a matched Python/NIXL throughput A/B for the terminal pipeline-parallel rendezvous derived from the production implementation in PR #16645.
Exact frozen candidate head:
fb3981ffa4d8e3e27f3a3fdc586198f359d4c5c2(tree976b1ae763b8111faebe2ace6b92c7a806b3883d). Its parent is the preserved failed experiment heada837ed1f49c7c154746b3c2f8347fedfcb957704. This TEST ONLY draft must never merge.The first failed paired head and the second invalid split head are preserved on backup branches. Their sanitized evidence is public:
Why the experiment was repaired
The earlier valid async-only run completed 256/256 requests at 454.92 output tok/s, but it had no matched synchronous control and therefore proved functionality rather than throughput uplift.
The first paired attempt used sequential arms in one allocation. Two requests hit the 600-second KV-transfer timeout, followed by backend failure and unbounded cleanup. One allocation also cannot safely contain both this roughly 77-minute workload and a potentially slower control under the four-hour allocation limit.
The second attempt split the arms into independent allocations, but it found strict harness defects before a causal comparison could be accepted: the nested arm environment did not reach outer CTX ranks, the lifecycle barrier counted GPU workers rather than outer pytest controllers, cumulative CTX evidence mixed profiling and measured runtimes, and allocation-specific request IDs made cross-stage digests incomparable. Both raw workloads completed, but both ran
terminal=0; their raw throughput values are intentionally rejected.Matched split design
One exact-head CI invocation requests two stage suffixes. Each stage uses a separate fresh three-node allocation built from the same downstream image:
Post-Merge-1: asynchronous terminal agreement (terminal=1,peer-ready=1).Post-Merge-2: legacy blocking terminal agreement (terminal=0,peer-ready=1).The two arms hold constant the exact source head, rendered image, model, selector, topology,
generation_first, Python/NIXL transfer, CTX PP4, GEN DEP8, metadata capacity 256, compute cap 4, 256 requests at concurrency 256, one round, dataset order, and instrumentation. Only terminal agreement changes.The repaired harness now:
sha256-length-prefixed-prompt-sha256-v1in evidence validation.Each arm retains bounded client, KV-transfer, coordination, and test deadlines sized above the observed healthy workload duration. Atomic failure propagation and bounded terminate/kill handling prevent a backend failure from leaving clients or barriers alive.
Acceptance criteria
Each arm must independently provide:
The asynchronous arm must additionally report 256 terminal commits on every CTX rank. Because the arms are separate stages, stage success alone is insufficient: the two prompt-content activation digests must be compared across the published stage artifacts before accepting the throughput delta. Missing or unequal evidence censors the comparison.
This experiment measures only the terminal-rendezvous effect. Because the arms use separate allocations, a single-run delta below roughly 3–5% is treated as inconclusive allocation noise.
Preflight
No replacement CI trigger will be posted until the automatic PR gates pass on the exact frozen head.