Skip to content

[NVBUG-6448152][test] TEST ONLY; DO NOT REVIEW: benchmark Python generation-first flag-off control - #16648

Closed
chienchunhung wants to merge 6 commits into
NVIDIA:mainfrom
chienchunhung:codex/python-consensus-gf-legacy
Closed

[NVBUG-6448152][test] TEST ONLY; DO NOT REVIEW: benchmark Python generation-first flag-off control#16648
chienchunhung wants to merge 6 commits into
NVIDIA:mainfrom
chienchunhung:codex/python-consensus-gf-legacy

Conversation

@chienchunhung

@chienchunhung chienchunhung commented Jul 20, 2026

Copy link
Copy Markdown
Collaborator

Purpose

TEST ONLY control for the generation-first Python-transceiver consensus performance experiment. The branch is exactly one signed YAML-only child of corrected harness base da6e140bda4c831bc813b3b192d7cb98a37025b2. That base is one harness-only child of the official asynchronous Python transceiver draft at product head 0f9acdd4dddbde7ccbf31973ce420bf4da6cb28a: it only passes the allowlisted server_config_extra.schedule_style value into the generated runtime config and adds focused tests. It does not change product behavior.

The previous run is not generation-first evidence: although its YAML requested generation_first, the old harness omitted that key from the generated server config and the runtime used context_first. This corrected head requires both the generated config and runtime log to confirm generation_first.

Experiment

Setting Value
Schedule generation_first
Async terminal agreement off
Async readiness agreement off
Context transceiver Python
Generation transceiver Python
Transport NIXL

The control explicitly sets both agreement flags to zero on the CTX workers, preventing ambient opt-in. The GEN workers use the Python transceiver but receive no CTX-only agreement flags.

The run must use the exact GB300 disaggregated perf-sanity selector, complete every submitted request, and publish an official output-token-throughput metric. Any failed request censors the throughput result.

This diagnostic is not intended to merge.

@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60498 [ run ] triggered by Bot. Commit: 93d1f7e Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60498 [ run ] completed with state FAILURE. Commit: 93d1f7e
/LLM/main/L0_MergeRequest_PR pipeline #48822 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chienchunhung
chienchunhung force-pushed the codex/python-consensus-gf-legacy branch from 93d1f7e to d2b0c45 Compare July 21, 2026 00:42
@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60513 [ run ] triggered by Bot. Commit: d2b0c45 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60513 [ run ] completed with state FAILURE. Commit: d2b0c45
/LLM/main/L0_MergeRequest_PR pipeline #48836 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chienchunhung chienchunhung changed the title [NVBUG-6448152][test] benchmark Python generation-first flag-off control [NVBUG-6448152][test] TEST ONLY; DO NOT REVIEW: benchmark Python generation-first flag-off control Jul 21, 2026
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
@chienchunhung
chienchunhung force-pushed the codex/python-consensus-gf-legacy branch from d2b0c45 to 11eb420 Compare July 21, 2026 23:33
@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60820 [ run ] triggered by Bot. Commit: 11eb420 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60820 [ run ] completed with state FAILURE. Commit: 11eb420
/LLM/main/L0_MergeRequest_PR pipeline #49095 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
…tment

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
@chienchunhung
chienchunhung force-pushed the codex/python-consensus-gf-legacy branch from 11eb420 to 156ca8e Compare July 22, 2026 03:38

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60874 [ run ] triggered by Bot. Commit: 156ca8e Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60874 [ run ] completed with state FAILURE. Commit: 156ca8e
/LLM/main/L0_MergeRequest_PR pipeline #49146 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants