[AutoDiff] Autodiff 12: Heap-backed adstack on LLVM backends (CPU/CUDA/AMDGPU)#537
Conversation
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a619832bc4
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
2c96718 to
8c2cc51
Compare
a619832 to
cd038de
Compare
8c2cc51 to
0bae8ab
Compare
cd038de to
04d17cd
Compare
There was a problem hiding this comment.
Additional findings (outside current diff — PR may have been updated during review):
-
🔴
quadrants/program/compile_config.h:54-68— The comment added to compile_config.h falsely claims SPIR-V now uses heap-backed StorageBuffers, but SPIR-V heap-backing is explicitly deferred to PR #493; SPIR-V still allocates adstacks as Function-scope per-thread on-chip memory bounded by the driver shader-compiler limit. The 32 to 256 default increase therefore multiplies per-thread private memory 8x on Metal/Vulkan, which can cause pipeline-creation failures for kernels that previously compiled correctly. The companion docs change compounds this by removing the only warning about the SPIR-V on-chip cap and replacing it with advice to bump default_ad_stack_size, which on SPIR-V causes shader-compiler rejection rather than heap growth.Extended reasoning...
What is wrong
compile_config.h lines 57-64 add a comment justifying the 32 to 256 default increase: "Both backends now heap-back the primal/adjoint slots: SPIR-V uses per-dispatch StorageBuffers (BufferType::AdStackHeapFloat + AdStackHeapInt, sliced by invocation)". This claim is factually incorrect for the current PR. The PR description itself explicitly states "SPIR-V side is the following PR", meaning PR #493 has not landed yet. After this PR merges, SPIR-V (Metal and Vulkan) still allocates every AdStackAllocaStmt using ir_->alloca_variable(arr_type) with spv::StorageClassFunction - Function-scope per-thread on-chip private memory - exactly as before.
Concrete code path
spirv_codegen.cpp:2221-2223 (unchanged in this diff) visits AdStackAllocaStmt and calls alloca_variable() for count_var, primal_arr, and adjoint_arr. There is no BufferType::AdStackHeapFloat or BufferType::AdStackHeapInt anywhere in the SPIR-V codegen, and no SPIR-V files appear in the list of changed files. The 8x raise is therefore applied to the SPIR-V path unconditionally by ControlFlowGraph::determine_ad_stack_size() in transforms/determine_ad_stack_size.cpp, which is arch-agnostic and falls back to default_ad_stack_size for any stack whose worst-case trip count cannot be statically proven.
Why existing code does not prevent it
The SPIR-V codegen has no guard that caps AdStackAllocaStmt::max_size against a per-thread on-chip budget - that responsibility fell on the deliberately-conservative 32-slot default. Removing that conservatism by raising the default 8x while the SPIR-V heap path is absent eliminates the only protection.
Impact
For a kernel with 4 f32 loop-carried variables and a dynamic loop whose trip count the compiler cannot prove:
- Old default (32): per-thread Function-scope demand = 4 * (8 + 3224) bytes ~= 1 KB
- New default (256): per-thread Function-scope demand = 4 * (8 + 25624) bytes ~= 8 KB
Apple Metal's MSL shader compiler rejects pipelines whose per-thread private-variable footprint exceeds its budget (typically 4-16 KB on consumer hardware), and the existing test test_adstack_shader_compile_failure_raises with ad_stack_size=65536 confirms this failure mode exists. A kernel that compiled and ran correctly at the 32-slot default may now fail at kernel-launch time with a pipeline-creation error whose message does not mention adstack size.
Step-by-step proof
- User calls qd.init(arch=qd.metal, ad_stack_experimental_enabled=True) - default_ad_stack_size is now 256 after this PR.
- A kernel has 4 f32 loop-carried variables under a range(n[None]) whose bound comes from a field.
- determine_ad_stack_size cannot prove the worst-case trip count, sets max_size = 256 for each adstack.
- SPIR-V codegen emits 4 x OpVariable StorageClassFunction arrays of 8x256 bytes each = 8192 bytes of per-thread on-chip private memory.
- Metal MSL compiler rejects the pipeline; compute.grad() raises RuntimeError: Failed to create pipeline with no hint that adstack sizing is the cause.
- The user consults the updated docs, which now say to bump default_ad_stack_size when they see a SPIR-V overflow. Doing so makes the pipeline failure worse, not better.
Docs regression
The removed Known Limitations bullet was the only explicit warning that on SPIR-V backends the adstack is allocated as per-thread on-chip memory capped by the driver. The new Tuning the capacity section and the Memory cost statement that "The buffer grows on demand to match the largest size any launch has needed so far" are presented without qualification and are false for SPIR-V - on SPIR-V the capacity is baked into the compiled shader at compile time, there is no grow-on-demand behavior, and bumping the value risks compile-time pipeline rejection rather than graceful heap growth.
Fix
Keep default_ad_stack_size at 32 until PR #493 lands. Update the comment in compile_config.h to remove the false SPIR-V StorageBuffer claim. Restore the Known Limitations bullet about the SPIR-V on-chip cap so Metal/Vulkan users receive the correct guidance.
0bae8ab to
eed8931
Compare
04d17cd to
4f43e77
Compare
eed8931 to
3d25bd7
Compare
4f43e77 to
10e5547
Compare
3d25bd7 to
c233cfb
Compare
10e5547 to
35b25a4
Compare
c233cfb to
e6ed50d
Compare
35b25a4 to
3a3e58c
Compare
e6ed50d to
d9c0752
Compare
3a3e58c to
98f2246
Compare
d9c0752 to
1240009
Compare
98f2246 to
c625fc5
Compare
|
|
||
| **Tuning the capacity.** Two `qd.init()` knobs control adstack sizing: | ||
|
|
||
| - `default_ad_stack_size=N` (default `256`): the fallback capacity for loops whose trip count the compiler cannot prove statically. Every adstack whose max_size was not deducible shares this value. Prefer tuning this knob, since it only affects the branch where the compiler needed to guess. |
There was a problem hiding this comment.
lets have the units please
default_ad_stack_size_mb
There was a problem hiding this comment.
if this is in units, or int32s, or something then maybe something like default_ad_stack_size_count or default_ad_stack_size_units or default_ad_stack_size_i32s`?
There was a problem hiding this comment.
what happens if there is an i64 in the loop?
| **Tuning the capacity.** Two `qd.init()` knobs control adstack sizing: | ||
|
|
||
| - `default_ad_stack_size=N` (default `256`): the fallback capacity for loops whose trip count the compiler cannot prove statically. Every adstack whose max_size was not deducible shares this value. Prefer tuning this knob, since it only affects the branch where the compiler needed to guess. | ||
| - `ad_stack_size=N` (default `0 = adaptive`): a hard override that forces every adstack in the program to exactly `N` slots, regardless of what the compiler proved. Prefer this knob only when a targeted experiment needs uniform sizing (e.g. stress-testing the runtime heap path). |
There was a problem hiding this comment.
units ad_stack_size_mb
| - `default_ad_stack_size=N` (default `256`): the fallback capacity for loops whose trip count the compiler cannot prove statically. Every adstack whose max_size was not deducible shares this value. Prefer tuning this knob, since it only affects the branch where the compiler needed to guess. | ||
| - `ad_stack_size=N` (default `0 = adaptive`): a hard override that forces every adstack in the program to exactly `N` slots, regardless of what the compiler proved. Prefer this knob only when a targeted experiment needs uniform sizing (e.g. stress-testing the runtime heap path). | ||
|
|
||
| **How to pick `default_ad_stack_size`.** The reverse pass of a `K`-iteration dynamic loop emits `K + 2` pushes per adstack (the trip count plus two setup pushes: one for the initial adjoint slot and one for the primal's starting value). Size the default at the flat trip count of the deepest unprovable dynamic loop in the program, plus that headroom. Common shapes: |
There was a problem hiding this comment.
default_ad_stack_size_mb
|
|
||
| **How to pick `default_ad_stack_size`.** The reverse pass of a `K`-iteration dynamic loop emits `K + 2` pushes per adstack (the trip count plus two setup pushes: one for the initial adjoint slot and one for the primal's starting value). Size the default at the flat trip count of the deepest unprovable dynamic loop in the program, plus that headroom. Common shapes: | ||
|
|
||
| - A single `qd.ndrange(n, m)` whose bounds come from a field: worst case is `n_max * m_max` iterations. Pick `N >= n_max * m_max + 2`. At `max_n_dofs_per_entity = 16`, 16 x 16 = 256 hits the default exactly. |
There was a problem hiding this comment.
dont we need to multiply by 4?
| **How to pick `default_ad_stack_size`.** The reverse pass of a `K`-iteration dynamic loop emits `K + 2` pushes per adstack (the trip count plus two setup pushes: one for the initial adjoint slot and one for the primal's starting value). Size the default at the flat trip count of the deepest unprovable dynamic loop in the program, plus that headroom. Common shapes: | ||
|
|
||
| - A single `qd.ndrange(n, m)` whose bounds come from a field: worst case is `n_max * m_max` iterations. Pick `N >= n_max * m_max + 2`. At `max_n_dofs_per_entity = 16`, 16 x 16 = 256 hits the default exactly. | ||
| - Nested `for i in range(a[None]): for j in range(b[None]):`: worst case is `a_max * b_max`, same rule. |
|
|
||
| - A single `qd.ndrange(n, m)` whose bounds come from a field: worst case is `n_max * m_max` iterations. Pick `N >= n_max * m_max + 2`. At `max_n_dofs_per_entity = 16`, 16 x 16 = 256 hits the default exactly. | ||
| - Nested `for i in range(a[None]): for j in range(b[None]):`: worst case is `a_max * b_max`, same rule. | ||
| - A single dynamic `for i in range(a[None])`: `N >= a_max + 2`. |
| - Nested `for i in range(a[None]): for j in range(b[None]):`: worst case is `a_max * b_max`, same rule. | ||
| - A single dynamic `for i in range(a[None])`: `N >= a_max + 2`. | ||
|
|
||
| **Memory cost.** The adstack pipeline allocates one small scratch buffer per loop-carried variable that the reverse pass has to remember. For example, a kernel whose dynamic loop reads and updates one float accumulator needs 1 adstack; a kernel whose loop updates four different floats needs 4. Integer counters and boolean branch flags used by the reverse pass also count (typically one each per dynamic `if` or nested loop). The total memory Quadrants allocates across all those buffers is roughly |
There was a problem hiding this comment.
why do we need to qualify with "that the reverse pass has to remember."? Are there loop-carried variables tha the reverse pass does not have to remember?
| **Memory cost.** The adstack pipeline allocates one small scratch buffer per loop-carried variable that the reverse pass has to remember. For example, a kernel whose dynamic loop reads and updates one float accumulator needs 1 adstack; a kernel whose loop updates four different floats needs 4. Integer counters and boolean branch flags used by the reverse pass also count (typically one each per dynamic `if` or nested loop). The total memory Quadrants allocates across all those buffers is roughly | ||
|
|
||
| ``` | ||
| num_threads * stack_size * bytes_per_element * num_loop_carried_variables |
There was a problem hiding this comment.
can you clarify where num_threads suddenly springs from? I'm guessing it's from the top level for loop, but you don't introduce t his I think. Or at least, I dont remember your introducing this.
| num_threads * stack_size * bytes_per_element * num_loop_carried_variables | ||
| ``` | ||
|
|
||
| where `bytes_per_element` depends on the element type and the backend. On the LLVM backends (CPU / CUDA / AMDGPU) each adstack slot stores both a primal and an adjoint value, so f32 costs 8, i32 costs 8, and bool costs 2 bytes per slot. On the SPIR-V backends (Metal / Vulkan) integer adstacks only store the primal (the reverse pass does not accumulate integer adjoints), and bool is widened to i32 at storage time because SPIR-V has no defined layout for `OpTypeBool`, so f32 costs 8, i32 costs 4, and bool costs 4 bytes per slot. The buffer lives on the device on GPU and in host RAM on CPU. `num_threads` is the number of threads the kernel actually dispatches, not a worst-case grid: on CPU this is the thread pool size (tens of threads), so the memory footprint stays small; on GPU it is the dispatched ndrange. The buffer grows on demand to match the largest size any launch has needed so far and is then reused across subsequent launches, so you do not need to reserve memory up front. |
There was a problem hiding this comment.
lets ditch the where, otherwsie no room to breathe. Seems lik a bunch new concepts here, so lets give the reader time to breathe. its a new paragraph.
There was a problem hiding this comment.
this all seesm like way too much detail. Do we really need to know this to use autodiff? Move it to an 'advanced' or 'under the hood' section.
There was a problem hiding this comment.
keep 'The buffer grows on demand to match the largest size any launch has needed so far and is then reused across subsequent launches, so you do not need to reserve memory up front.'
| **Problem.** Reverse-mode AD through a dynamic loop (one whose trip count is not known at compile time) needs to recover the primal value at each iteration when walking the loop backwards. Without that, the chain-rule steps read a stale value and the gradients come out silently wrong. Static-unrolled (`qd.static(range(...))`) loops are not affected because every iteration becomes its own inlined block at compile time. | ||
|
|
||
| **How Quadrants does it.** An opt-in compiler pipeline called the *autodiff stack* (*adstack*) allocates a per-variable stack alongside each loop-carried primal. The forward pass pushes an entry each iteration; the reverse pass pops them back off in reverse order to recover the correct primal for every chain-rule step. It is opt-in because it costs extra per-thread memory and compile time, and because most kernels do not need it. Running with adstack enabled when it is not strictly needed is safe. Running without it when it is needed raises a `QuadrantsCompilationError` in most cases (the autodiff pass rejects a non-static range that would otherwise lose its primal); in the narrow cases where the kernel compiles anyway, the reverse pass reads a stale value for every iteration and the gradients come out wrong but non-zero. | ||
| **How Quadrants does it.** An opt-in compiler pipeline called the *autodiff stack* (*adstack*) allocates a per-variable stack alongside each primal that is updated inside the loop and therefore changes from one iteration to the next. The forward pass pushes an entry each iteration; the reverse pass pops them back off in reverse order to recover the correct primal for every chain-rule step. It is opt-in because it costs extra per-thread memory and compile time, and because most kernels do not need it. Running with adstack enabled when it is not strictly needed is safe. Running without it when it is needed raises a `QuadrantsCompilationError` in most cases (the autodiff pass rejects a non-static range that would otherwise lose its primal); in the narrow cases where the kernel compiles anyway, the reverse pass reads a stale value for every iteration and the gradients come out wrong but non-zero. |
There was a problem hiding this comment.
break the first sentence into two. so each setnence just states a single concept.
There was a problem hiding this comment.
" the autodiff stack (adstack) " => "adstack". We only ever refer to it as adstack, so let's say adstack is its name. We can put the long form in brakcets i youf want "called the adstack (short for "(a)uto(d)iff (stack)")"
There was a problem hiding this comment.
I think remove "and therefore changes from one iteration to the next"
There was a problem hiding this comment.
" It is opt-in" => "adstack is opt-in"
There was a problem hiding this comment.
"Running without it when it is needed raises a QuadrantsCompilationError in most cases (the autodiff pass rejects a non-static range that would otherwise lose its primal);" => nice 🙌
There was a problem hiding this comment.
"in the narrow cases where the kernel compiles anyway, the reverse pass reads a stale value for every iteration and the gradients come out wrong but non-zero." => would be nice to get rid of such exceptional cases. Do we know what they are? Can we document them?
There was a problem hiding this comment.
I know it's not part of your changes in this pr, but it's part of your set of prs overall, and anyway, I think it would be good to address:
I couldnt understand this sentence. I mean, I probably could if I thought about it a long time, bu I think it would be good to explain step by step, so people can just read fluidly, and undersatnd it, without grinding to a halt, and having to work through stuff in their head.
| - A loop-carried dependency (a variable read, written, and read again across iterations, e.g. `v = v * 0.95 + 0.01`). | ||
| - A loop-carried variable - one whose value is carried forward from each iteration into the next, e.g. `v = v * 0.95 + 0.01`. | ||
| - A local variable used as an index into a global field. | ||
| - Non-linear ops (`sin`, `cos`, `exp`, `sqrt`, `tanh`, `pow`, ...) whose derivative depends on the primal value at that iteration. |
There was a problem hiding this comment.
can we have a counter-example of a non-linear op that doesnt need adstack, in the 'do not need it' section below please.
| | nested `for i in range(a[None]): for j in range(b[None])` | `a_max * b_max + 2` | | ||
| | `qd.ndrange(n, m)` with field-derived `n`, `m` | `n_max * m_max + 2` | | ||
|
|
||
| At `max_n_dofs_per_entity = 16`, a 16 x 16 ndrange hits the default exactly (`256`). |
There was a problem hiding this comment.
mixtue of fonts is visually distracting. either put backticks around 16 x 16 too, or remove from = 16
|
|
||
| At `max_n_dofs_per_entity = 16`, a 16 x 16 ndrange hits the default exactly (`256`). | ||
|
|
||
| **Memory footprint.** The pipeline allocates one scratch buffer per piece of reverse-pass state. That count includes every loop-carried variable the reverse pass has to replay, plus any integer counter and any boolean branch flag it has to read back. Total memory across all buffers is approximately |
There was a problem hiding this comment.
whats a 'piece'? per stack carried variable?
There was a problem hiding this comment.
oh you define it next 🤔
There was a problem hiding this comment.
Maybe we just avoid the issue by avoiding having to use this noun at all? for example
"The pipeline allocates a scratch buffer for each loop-carried variable, and also for any loop counter, and any boolean branch flags." ?
Still, I'm unclear about this 'integer counter' and 'boolean branch flag'. You havent defined them before. Could you define these in a previous paragraph please.
|
|
||
| `num_threads` is the number of threads the kernel actually dispatches. On CPU that is the thread-pool size, typically tens. On GPU it is the full ndrange. `bytes_per_slot` scales with the element's storage size and the backend; see the two tables below. | ||
|
|
||
| On LLVM backends (CPU / CUDA / AMDGPU), each adstack slot stores both a primal and an adjoint value, so `bytes_per_slot = 2 * sizeof(T)` for every element type `T`. Common cases: |
| At `max_n_dofs_per_entity = 16`, a `16 x 16` ndrange hits the default exactly (`256`). | ||
|
|
||
| **Memory footprint.** The pipeline allocates one scratch buffer per piece of reverse-pass state. That count includes every loop-carried variable the reverse pass has to replay, plus any integer counter and any boolean branch flag it has to read back. Total memory across all buffers is approximately | ||
| **Memory footprint.** With one scratch buffer per adstack (see above), the total memory cost depends on two further quantities. The first is the number of threads the kernel actually dispatches, which we call `num_threads`. On CPU that is the thread-pool size, typically tens. On GPU it is the full ndrange. The second is `bytes_per_slot`, which scales with the element's storage size and the backend; the two tables below work through its concrete values. Total memory across all buffers is then approximately: |
There was a problem hiding this comment.
have we used the term 'element' thus far? Have we defined it?
|
Checklist:
=> ok to merge |
7c71e52 to
35ff6a8
Compare
ae97db1 to
92198fa
Compare
* [Misc] Warn user to disable caching when print_ir/QD_DUMP_IR enabled (Genesis-Embodied-AI#425) Co-authored-by: v01dxyz <v01dxyz@v01d.xyz> * [Build] Pin torch version to CUDA 12.8 for CUDA tests (Genesis-Embodied-AI#428) * [Misc] Fixing up taichi-dev urls (Genesis-Embodied-AI#429) * [Perf] Rename cuda_graph to gpu_graph across the codebase (Genesis-Embodied-AI#430) * Misc: fix typo integeral -> integral (Genesis-Embodied-AI#434) Co-authored-by: v01dxyz <v01dxyz@v01d.xyz> * [Perf] CUDA graph 4: call from multiple locations (Genesis-Embodied-AI#420) * [Bug] Fix fastcache not restoring graph_do_while_arg (Genesis-Embodied-AI#435) * [Perf] Cache last-call result in perf_dispatch for single-compatible case (Genesis-Embodied-AI#438) * Fix gpu_graph fallback on old Nvidia GPU. (Genesis-Embodied-AI#443) * Fix shared memory offset not reset between CUDA kernels. (Genesis-Embodied-AI#442) * [Misc] Allow disabling GPU graph via QD_GPU_GRAPH=0 env var (Genesis-Embodied-AI#439) * [Misc] Add named top-level loops (Genesis-Embodied-AI#440) * [Misc] Rename gpu_graph to graph (Genesis-Embodied-AI#446) * [Misc] Add cross-platform shuffle (Genesis-Embodied-AI#447) * [Bug] Fix graph_do_while on Windows: search for cudadevrt.lib (Genesis-Embodied-AI#456) * [Bug] Also search default CUDA toolkit install location on Windows (Genesis-Embodied-AI#461) * [SPIRV] Feature Parity Atomics & Shared Array (Genesis-Embodied-AI#432) * [Misc] Change clang format to 120 characters (Genesis-Embodied-AI#463) * [Misc] CUDA graph 5 Add fatbin (Genesis-Embodied-AI#464) * [Bug] Reuse VkInstance across init/reset cycles (Genesis-Embodied-AI#465) * [Perf] Tiles 1: _load, _store, _eye_ (Genesis-Embodied-AI#466) * [Misc] Remove dead InternalFuncStmt type_check override (Genesis-Embodied-AI#471) * [Perf] Tiles 2: add cholesky and ger (Genesis-Embodied-AI#472) * [Perf] Tiles 2b: add triangular solve (Genesis-Embodied-AI#474) * [Misc] Refactor: use _get_col/_set_col in tiles load/store/init (Genesis-Embodied-AI#475) * [Build] Fix flaky test_clock_accuracy (Genesis-Embodied-AI#436) * Fix AARCH64 emitting invalid asm in CUDA kernels. (Genesis-Embodied-AI#473) Co-authored-by: Hugh Perkins <hughperkins@gmail.com> * [AMDGPU] Enable HIP memory pool and surface pool-exhaustion errors. (Genesis-Embodied-AI#485) * [AMDGPU] Scope hsaco tmp dir per-user to avoid collisions. (Genesis-Embodied-AI#484) * [Perf] Tiles 3: Add slice syntax, qd.outer() and initial doc (Genesis-Embodied-AI#477) * [AMDGPU] Fix gradient computation. (Genesis-Embodied-AI#486) * Enable all backends that are supported in unit tests. (Genesis-Embodied-AI#488) * Fix SPIRV ID overflow for large kernels due to autodiff. (Genesis-Embodied-AI#489) * [Misc] Fix purity checker to allow accessing constants from quadrants modules (Genesis-Embodied-AI#487) * [Misc] Increase tolerance for clock monotonic test (Genesis-Embodied-AI#492) * [CI] Serialize api doc workflow (Genesis-Embodied-AI#494) * [CI] Increase tolerance for clock test (Genesis-Embodied-AI#506) * [CI] Increase clock test tolerance to 20% (Genesis-Embodied-AI#509) * [Perf] Add tensor_type parametrization to tile16 tests (Genesis-Embodied-AI#504) * [Perf] Tiles 4b: Migrate tiles16 tests to enable fastcache (Genesis-Embodied-AI#505) * [Perf] Tiles 4c: add Tiles16x16 proxy (Genesis-Embodied-AI#507) * [Perf] Tiles 4d: Consolidate slice error tests using parametrize (Genesis-Embodied-AI#508) * [Perf] Tiles 4: add SharedArray slice support (Genesis-Embodied-AI#482) * [Perf] Tiles 5: add Cholesky benchmark demo (Genesis-Embodied-AI#483) * [Doc] Add user guide page for subgroup shuffle (Genesis-Embodied-AI#512) * [Perf] Implement cross-platform shuffle_down (Genesis-Embodied-AI#510) * [Perf] Add portable subgroup reduce_add and reduce_all_add (Genesis-Embodied-AI#511) * [Perf] Add first warmup config to perf dispatch (Genesis-Embodied-AI#422) * [AutoDiff] Autodiff 1: Add baseline adstack regression test for unary_collections (Genesis-Embodied-AI#500) * [AutoDiff] Autodiff 2: Implement derivative for tan (Genesis-Embodied-AI#501) * [AutoDiff] Autodiff 3: Recompute tanh/exp on the operand in the reverse pass (Genesis-Embodied-AI#502) * [AutoDiff] Autodiff 4: Mark rsqrt as non-linear for adstack promotion (Genesis-Embodied-AI#503) * [AutoDiff] Autodiff 5: Fix adjoint-alloca placement for GlobalLoads outside the current range-for (Genesis-Embodied-AI#496) * [AutoDiff] Autodiff 6: Adstack regression tests (Genesis-Embodied-AI#491) * [AutoDiff] Autodiff 7: Fix header size in AdStackAllocaStmt to match u64 runtime layout (Genesis-Embodied-AI#534) * [AutoDiff] Autodiff 8: Surface LLVM adstack push/pop overflow as a Python exception (Genesis-Embodied-AI#535) * [AutoDiff] Autodiff 9: Guard against LLVM worker-thread stack overflow from large per-task adstack budget (Genesis-Embodied-AI#495) * [AutoDiff] Autodiff 10: Implement adstack for SPIR-V (Genesis-Embodied-AI#490) * [AutoDiff] Autodiff 11: Latent adstack-adjacent fixes (AMDGPU hipFree, flush() keeps ctx_buffers_, always-preallocate) (Genesis-Embodied-AI#536) * [Doc] Add AGENTS.md with instructions for AI agents (Genesis-Embodied-AI#541) * [Bug] Abort kernel execution on assertion failure instead of segfaulting (Genesis-Embodied-AI#419) * [Type] ndarray typing 1: Add eval_str=True to inspect.signature() calls (Genesis-Embodied-AI#411) * [CI] Suppress reportPrivateImportUsage in torch-using files (Genesis-Embodied-AI#552) * [Misc] QD_DUMP_IR dumps to files with the task_id added to the filename (Genesis-Embodied-AI#441) * [Type] ndarray typing 2: Fix NDArray single-arg subscript crash (Genesis-Embodied-AI#412) * [Test] Flush xdist channel before worker exit so test failure reports are visible (Genesis-Embodied-AI#555) * [CI] Reduce test retries on CI from 3 to 1. (Genesis-Embodied-AI#554) * [AutoDiff] Autodiff 12: Heap-backed adstack on LLVM backends (CPU/CUDA/AMDGPU) (Genesis-Embodied-AI#537) * [AutoDiff] Autodiff 13: Heap-backed adstack on SPIR-V backends (Metal, Vulkan) (Genesis-Embodied-AI#493) * [AutoDiff] Autodiff 14: Resolve bounded-inner-loop adstacks without default_ad_stack_size fallback (Genesis-Embodied-AI#539) * [SPIRV] Vulkan SPIR-V correctness: atomic-view aliasing, PSB stride, narrow storage caps, u1 cast, per-init layer recheck (Genesis-Embodied-AI#513) * [Build] Autodiff 15: Replace 2022 MoltenVK pin with LunarG Vulkan SDK fetch and sanitise MoltenVK cap advertisement (Genesis-Embodied-AI#551) * [Test] Suppress stock pytest-timeout to avoid conflict with pytest_hardtle (Genesis-Embodied-AI#557) * [Vulkan] Use SDK validation layer for debugPrintf instead of apt package (Genesis-Embodied-AI#562) * [Test] Fix flaky perf_dispatch tests by increasing work amounts (Genesis-Embodied-AI#559) * [Test] Add --maxfail CLI option to run_tests.py (default 20) (Genesis-Embodied-AI#558) * [CI] Vulkan debug printf fix to address flaky tests (Genesis-Embodied-AI#563) * [Docs] Add a new page to help for first time contributors (Genesis-Embodied-AI#426) Authored-by: v01dxyz <v01dxyz@v01d.xyz> * [AutoDiff] Autodiff 16: Resolve reverse-mode adstack depths per-launch via runtime-evaluated SizeExpr (Genesis-Embodied-AI#543) * Fix: raise error if device memory allocation fails (Genesis-Embodied-AI#451) (Genesis-Embodied-AI#453) Co-authored-by: v01dxyz <v01dxyz@v01d.xyz> Co-authored-by: Hugh Perkins <hughperkins@gmail.com> * [CI] Add CI job to check line wrapping of comments and docs (Genesis-Embodied-AI#564) * [Misc] Add coverage report to PRs, including kernels (Genesis-Embodied-AI#470) * [CI] CI wrap check feeds only diffs to agent (Genesis-Embodied-AI#567) * Skip 'flaky' test on MacOS CI. (Genesis-Embodied-AI#573) * [Test] Fix missing `import sys` in test_fail_device_memory_allocation (Genesis-Embodied-AI#574) * [CI] Fix Vulkan debugPrintf flake with session-scoped warmup (Genesis-Embodied-AI#571) * [AutoDiff] determine_ad_stack_size: replace whole-CFG Bellman-Ford with SCC + DAG DP (Genesis-Embodied-AI#575) * [Test] Fix macOS OOM skip reason to describe actual root cause (Genesis-Embodied-AI#576) * [Lang] whole_kernel_cse: 2.5x compile time speedup on large kernels (Genesis-Embodied-AI#577) * [CI] Add CI check for unnecessarily deleted comments (Genesis-Embodied-AI#570) * [CI] Migrate coverage report to github Check page (Genesis-Embodied-AI#566) * [Lang] Skip IR verifier between passes unless debug=true (Genesis-Embodied-AI#579) * [Lang] Inline AdStack ops on release LLVM codegen: dramatically reduces compile time for adstack-enabled reverse-mode kernels (Genesis-Embodied-AI#584) * [CUDA] Honor offline_cache=False end-to-end so QD_OFFLINE_CACHE=0 actually gives a cold compile (Genesis-Embodied-AI#580) * [Type] Tensor 24 (Genesis-Embodied-AI#561) Co-authored-by: hugh <hugh@slurm-login-0.slurm-login.tenant-slurm.svc.cluster.local> * [Lang] auto_diff host-walk reductions: dramatically faster front-end compile time on adstack-enabled reverse-mode kernels (Genesis-Embodied-AI#587) * [AutoDiff] Speed up reverse-mode kernel launches on GPU backends (Genesis-Embodied-AI#578) * [Vulkan] Move adstack-sizer scratch out of Function-scope memory to fix SPIR-V pipeline build failures (Genesis-Embodied-AI#588) * [AutoDiff] Improve diagnosis of unsupported reverse-mode AD patterns (Genesis-Embodied-AI#590) * [Bug] Fix: promote Ndarray to AnyArray in build_Name for flattened struct fields (Genesis-Embodied-AI#592) * [SPIR-V] Shrink reverse-grad kernel MSL by ~50% (Genesis-Embodied-AI#591) * [CI] Add CI check that PR changes have test coverage (Genesis-Embodied-AI#596) * [Perf] Enable zero-copy in to_torch() and to_numpy() (Genesis-Embodied-AI#450) * Add BufferView: safe sub-range ndarray access for kernels (Genesis-Embodied-AI#585) Co-authored-by: alanray-tech <alanray-tech@users.noreply.github.com> Co-authored-by: Hugh Perkins <hughperkins@gmail.com> * [Doc] Add user-facing fastcache documentation (Genesis-Embodied-AI#597) Co-authored-by: hugh <hugh@slurm-login-0.slurm-login.tenant-slurm.svc.cluster.local> * [Misc] Upgrade to enable v1 dlpack so to_numpy(copy=False) writable (Genesis-Embodied-AI#598) Co-authored-by: root <root@rtx-209-201.slurm-compute.tenant-slurm.svc.cluster.local> * [AutoDiff] Cut reverse-mode adstack memory usage 10x on all backends (Genesis-Embodied-AI#599) * [Misc] Add CI check for feature file factorization (Genesis-Embodied-AI#606) * [Perf] Skip _recursive_set_args for all-Field frozen dataclass structs (Genesis-Embodied-AI#607) Co-authored-by: Cursor <cursoragent@cursor.com> * [AutoDiff] SNode-arm bound-expr capture rejects fold-attack gate indices (Genesis-Embodied-AI#610) * [Misc] Suppress field fastcache warning for qd.Tensor (Genesis-Embodied-AI#615) Co-authored-by: Cursor <cursoragent@cursor.com> * [AutoDiff] Adstack heap: clip reducer count by per-task loop trip count (compile-time and SizeExpr-evaluated) (Genesis-Embodied-AI#611) * [Misc] Forward copy= through qd.Tensor, add copy=None option (Genesis-Embodied-AI#616) Co-authored-by: Cursor <cursoragent@cursor.com> * [Doc] Update README (Genesis-Embodied-AI#617) Co-authored-by: Cursor <cursoragent@cursor.com> * [CI] Fix coverage report showing def lines as uncovered (Genesis-Embodied-AI#623) Co-authored-by: Cursor <cursoragent@cursor.com> * [Perf] Generic launcher: persistent context, JIT-pointer reuse, Metal compute encoder, LLVM-GPU async memory ops (Part 1/2) (Genesis-Embodied-AI#619) * [CI] Encode Python-first testing policy in coverage-check prompt (Genesis-Embodied-AI#622) Co-authored-by: Cursor <cursoragent@cursor.com> * [CI] Add PR Line change report (Genesis-Embodied-AI#624) Co-authored-by: Cursor <cursoragent@cursor.com> * [CI] Disable quadrants pytest plugin during quadrants internal coverage runs (Genesis-Embodied-AI#629) Co-authored-by: Cursor <cursoragent@cursor.com> * [AutoDiff] Adstack load+store eliminations: EliminateRecomputableAdStackPushes pass + leaf extensions (Genesis-Embodied-AI#621) * [CI] Simplify coverage PR comment to a single linked line (Genesis-Embodied-AI#630) * [CUDA] Add AGX Thor, SM_110 (Genesis-Embodied-AI#631) Co-authored-by: Johnny Nunez and Hugh Perkins * [CI] Lines changed report: collapse PR comment to a single linked totals line (Genesis-Embodied-AI#632) * [FEATURE] Support external Metal command queue via qd.init (Genesis-Embodied-AI#618) Co-authored-by: Cursor <cursoragent@cursor.com> * [Perf] Cache adstack-sizer metadata per task across SPIR-V + LLVM-GPU; per-snode / DeviceAllocation invalidation (Part 2/2) (Genesis-Embodied-AI#620) * [AutoDiff] Disable EliminateRecomputableAdStackPushes pending mutated-SNode chain-leaf fix (Genesis-Embodied-AI#633) * [AutoDiff] Adstack chain-clone safety: mutated-SNode leaf reject + load_top consumer-aware guard (Genesis-Embodied-AI#634) * [Docs] Add user-guide page for qd.simt.block.* primitives (Genesis-Embodied-AI#638) * [Docs] Expand qd.simt.subgroup user-guide page to cover every op (Genesis-Embodied-AI#639) * [Perf] Streams 1-4 (Genesis-Embodied-AI#410) * [Docs] Add user-guide page for matrix decompositions and solvers (Genesis-Embodied-AI#643) * [Bug] Revert "[Perf] Streams 1-4 (Genesis-Embodied-AI#410)" (Genesis-Embodied-AI#650) * [Docs] Add user-guide page for atomics and bit operations (Genesis-Embodied-AI#640) * [Docs] Add user-guide page for qd.simt.grid.* primitives (Genesis-Embodied-AI#641) * [AutoDiff] Adstack max-reducer: parallel multi-axis MaxOverRange dispatch (Genesis-Embodied-AI#635) * [AMDGPU] Fix amdgpu parallel rand init (Genesis-Embodied-AI#658) * [Perf] Adstack: skip max-reducer recognizer on CPU + lift host-eval cap (Genesis-Embodied-AI#655) * [Perf] Re-land Streams 1-4 with bug fixes (Genesis-Embodied-AI#653) * [AMDGPU] Apply device_memory_GB=0.3 cap to AMDGPU tests (Genesis-Embodied-AI#659) * [Perf] Per-launch host sync: drop wait_idle on SPIR-V, pin stream and drop stream_synchronize on CUDA/AMDGPU (Genesis-Embodied-AI#654) * [AMDGPU] Unload hipModule_t in JITModuleAMDGPU destructor (Genesis-Embodied-AI#660) * [AMDGPU] Trim default mempool on qd.reset() (Genesis-Embodied-AI#669) * [AMDGPU] Hoist rand-state buffer to process lifetime (Genesis-Embodied-AI#668) * [Streams] Use events for streams serialization on AMDGPU and CUDA (Genesis-Embodied-AI#667) * [Perf] Adstack max-reducer: launch cache + zero-copy result map; content-stable registry_id (Genesis-Embodied-AI#671) * [SPIR-V] dispatch_max_reducers: register each task with the real kernel name (Genesis-Embodied-AI#675) * [AutoDiff] Debug-mode field/grad/dual: dtype, layout, and access-time invariants (Genesis-Embodied-AI#677) * [Docs] Add user-guide page for qd.algorithms.* device-wide algorithms (Genesis-Embodied-AI#642) Co-authored-by: alanray-tech <alan.ray@genesis-ai.company> * [Docs] Doc for existing atomics: switch support table to per-backend columns (Genesis-Embodied-AI#657) Co-authored-by: alanray-tech <alan.ray@genesis-ai.company> * [GPU] Cross gpu atomics (Genesis-Embodied-AI#666) Co-authored-by: alanray-tech <alan.ray@genesis-ai.company> * [GPU] Make block operations portable cross-gpu (Genesis-Embodied-AI#664) * [Perf] CPU LLVM adstack-cache: skip per-launch bump-writes + ndarray_shapes capture on forward-only handles (Genesis-Embodied-AI#685) * [GPU] Cross-GPU for grid ops (Genesis-Embodied-AI#670) * [Math] Make bitop operations portable cross-gpu (Genesis-Embodied-AI#662) * [AMDGPU] Always use wave64, on both RDNA and CDNA (Genesis-Embodied-AI#687) * [AMDGPU] Use syncscope("agent") for atomix xor to avoid CAS livelock (Genesis-Embodied-AI#672) * [GPU] New bit ops for QIPC (Genesis-Embodied-AI#679) * [GPU] Subgroup ops cross-gpu (Genesis-Embodied-AI#665) * [Graph] Rename CUDA Graph to Graph in docs (Genesis-Embodied-AI#691) * [SPIR-V] Fix FIFO-queue ordering when sharing command queue. (Genesis-Embodied-AI#694) * [Atomics] New QIPC ops for atomics (Genesis-Embodied-AI#690) * Pass dataclass sub-structs into qd.func (Genesis-Embodied-AI#698) * [AMDGPU] HIP graph runtime support for @qd.kernel(graph=True) (Genesis-Embodied-AI#692) * [CI] Add per-file timing report to Mac Metal test job (Genesis-Embodied-AI#695) Co-authored-by: Cursor <cursoragent@cursor.com> * [CI] Enable kernel disk cache during tests (Genesis-Embodied-AI#696) * [Math] New QIPC ops for single-threaded linalg (Genesis-Embodied-AI#683) * [BREAKING][GPU] New QIPC ops for subgroups (Genesis-Embodied-AI#676) * [GPU] New QIPC ops for block (Genesis-Embodied-AI#684) * [GPU] New device-level ops for QIPC (Genesis-Embodied-AI#693) * [algorithms] PrefixSumExecutor: drop unused GRID_SZ local (Genesis-Embodied-AI#701) * [block] sync(): fix unsupported-arch error message (Genesis-Embodied-AI#700) * [volatile_load] add qd.volatile_load primitive (closes Genesis-Embodied-AI#648) (Genesis-Embodied-AI#702) * [AutoDiff] Reject recycled identity_key in AdStackCache::register_adstack_sizing_info (Genesis-Embodied-AI#708) * [Vulkan] Declare GroupNonUniform SPIR-V caps and enable shaderSubgroupExtendedTypes (Genesis-Embodied-AI#707) * Fix duplicate HIP graph driver-function declarations after v1.0.0 merge The amd-integration fork had cherry-picked the HIP graph driver functions (graph_create / graph_destroy / graph_add_kernel_node / graph_instantiate / graph_exec_destroy / graph_launch), and upstream v1.0.0 added the same set. The per-file 3-way merge appended both copies into amdgpu_driver_functions.inc.h, producing redeclaration errors that broke the AMDGPU RHI/runtime compile. Drop the upstream duplicate block; the signatures are identical to the fork's existing declarations. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix AMDGPU launcher coherence and num_instructions visibility after v1.0.0 merge - kernel_launcher.cpp: the 3-way merge spliced upstream v1.0.0's launch_llvm_kernel rewrite (ephemeral arg/context buffers, explicit-stream path, AmdgpuDefaultStream PinGuard) onto the AMD fork's kernarg-by-value + persistent-scratch design, leaving references to undefined `ephemeral_context_ptr`. Restore the fork's coherent launch_llvm_kernel verbatim; it calls the (already merged) enhanced launch_offloaded_tasks, which keeps the max-reducer dispatch and stream-parallel groups adapted onto the AMD launch path. - llvm_context.h: both the fork and upstream added `num_instructions`; the merge kept upstream's private placement, but the AMDGPU codegen force-inline heuristic calls it statically from outside the class. Move it back to the public section. Co-authored-by: Cursor <cursoragent@cursor.com> * Restore async result D2H and hoist kernarg vectors in AMDGPU launcher The v1.0.0 merge resolution regressed two amd-integration baseline optimizations in launch_llvm_kernel / launch_offloaded_tasks: - The per-launch result-buffer copy was a blocking memcpy_device_to_host, forcing a host stall on every value-returning launch and serializing the GPU pipeline. Restore the async D2H (the caller synchronizes lazily when it needs the value); external-array transfers still stream_synchronize once before reading back. - launch_task constructed the kernarg std::vectors from initializer lists ({kernarg_payload} / {kernarg_size}) on every dispatch (heap alloc + free per launch). Hoist arg_ptrs/arg_sizes out of the per-task launch and reuse. Co-authored-by: Cursor <cursoragent@cursor.com> * amdgpu: default to LDS permlane64 emulation; drop host-x86 barrier asm on retarget Two AMDGPU JIT-compile crashes surfaced after the v1.0.0 merge pulled in the QIPC subgroup ops (Genesis-Embodied-AI#676), which made the rigid constraint solver's wave-cooperative reductions route through `amdgpu_cross_half_shuffle_i32`. Both manifested as a SIGSEGV inside `llvm::SIInstrInfo::getInstSizeInBytes` during `JITSessionAMDGPU::compile_module_to_hsaco` (i.e. at first kernel launch), and reproduce on gfx942 / MI300X. Baseline 0.4.6 never emitted these constructs, which is why it was unaffected. 1. Native `llvm.amdgcn.permlane64` lowering crashes the bundled LLVM 22.1.0 AMDGPU backend. Default `amdgpu_permlane64` to the existing LDS-roundtrip software emulation on every target (it produces identical results). Add `QD_AMDGPU_USE_NATIVE_PERMLANE64=1` to opt back into the native instruction once the backend bug is fixed; the old `QD_AMDGPU_FORCE_PERMLANE64_FALLBACK` is now the default and still honored. This is the actual crash fix. 2. The runtime module is compiled by the host x86_64 clang and only retargeted to amdgcn here, so `amdgpu_cross_half_shuffle_i32`'s `__asm__ volatile("" : "+v"(byte))` optimization barrier carries x86 flag clobbers (`~{dirflag},~{fpsr},~{flags}`) that are meaningless on AMDGPU. The IR verifies but the empty-body INLINEASM is invalid on the amdgcn target. Neutralize empty-body barrier asm during retarget (forward the tied value, then erase) so no stale host asm reaches codegen. On the wave64 targets we ship `ds_bpermute` already addresses the full wave, so the hint is a no-op. Co-authored-by: Cursor <cursoragent@cursor.com> * style: apply clang-format (v19.1.7) to AMDGPU fn_attrs and launcher sources CI pre-commit's clang-format hook reformatted these files (long declarations/lambda signatures collapsed onto single lines per the repo's clang-format config). Apply the same formatting so the hook passes. No functional changes. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(amdgpu): use CreateNeg for branchless i32 sgn instead of CreateSub(0, input) clang-tidy (modernize-use-nullptr, -warnings-as-errors) flagged `builder->CreateSub(0, input)` in the i32 sgn path: the literal `0` binds to the `llvm::Value*` LHS parameter as a null pointer, not an integer zero. Replace with `builder->CreateNeg(input)`, which emits `0 - input` with a proper zero constant -- identical intended semantics, and clang-tidy clean. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Robert Dazi <14996868+v01dXYZ@users.noreply.github.com> Co-authored-by: v01dxyz <v01dxyz@v01d.xyz> Co-authored-by: Hugh Perkins <hughperkins@gmail.com> Co-authored-by: Alexis DUBURCQ <alexis.duburcq@gmail.com> Co-authored-by: hugh <hugh@slurm-login-0.slurm-login.tenant-slurm.svc.cluster.local> Co-authored-by: alanray-tech <alan.ray@genesis-ai.company> Co-authored-by: alanray-tech <alanray-tech@users.noreply.github.com> Co-authored-by: root <root@rtx-209-201.slurm-compute.tenant-slurm.svc.cluster.local> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Johnny <johnnynuca14@gmail.com>
Heap-backed adstack on LLVM backends (CPU / CUDA / AMDGPU)
TL;DR
Prior behaviour on CPU / CUDA / AMDGPU: every
AdStackAllocaStmtlowered to a function-scopeallocaat the task's entry block, so every adstack lived on the LLVM stack frame (= worker-thread stack on CPU, per-thread local memory on GPU). A kernel with many loop-carried values atdefault_ad_stack_size=256crossed the worker-thread limit and silently corrupted adjacent stack memory; the previous PR in the stack added a 256 KB codegen-time guard that hard-aborted those kernels.This PR moves the storage off the stack:
{ad_stack_offsets_, ad_stack_per_thread_stride_}, andvisit(AdStackAllocaStmt)emitsbase = runtime->adstack_heap_buffer + linear_thread_idx * stride + offsetinstead of an alloca. Base is loaded once inentry_blockand reused.LlvmRuntimeExecutor::ensure_adstack_heap(needed_bytes)grows the per-runtime slab via amortised doubling, publishes the new pointer/size intoruntime->{adstack_heap_buffer, adstack_heap_size}by caching the two device field-pointer addresses on first grow and writing through them on every subsequent grow.ensure_adstack_heap(per_thread_stride * num_threads)before each task launch. Dynamic-bound range-for tasks resolvenum_threadsby readingbegin/endfromruntime->temporariesvia a host-side DtoH memcpy.per_thread_stride > 0, because graph baking precludes the host-sideensure_adstack_heapstep between dispatches.default_ad_stack_sizeexposed viaqd.init(); raised from 32 → 256 now that the per-thread on-chip / worker-stack budget no longer caps it.Nothing changes for kernels that don't enable the adstack extension.
Why
Prior to this PR, a kernel like a reverse-mode articulated-body dynamics step in Genesis hit the 256 KB CPU-stack budget at modest capacities (4 loop-carried f64 variables × 4096 entries × 16 bytes each already crosses it). The two alternatives — ship with
default_ad_stack_sizecapped at a value small enough to fit on every worker stack, or ask users to lowerad_stack_sizeper-kernel — either regress correctness on large kernels or force tuning noise on the user. Moving the storage off-stack removes the constraint entirely: per-thread slice size is bounded only bynum_threads * per_thread_strideand the driver's allocator.Changes
Codegen (
quadrants/codegen/llvm/codegen_llvm.{h,cpp},llvm_compiled_data.h)TaskCodeGenLLVMgrows three new per-task fields:ad_stack_per_thread_stride_— sum ofAdStackAllocaStmt::size_in_bytes()(aligned up to 8) for every adstack in the task.ad_stack_offsets_— map from each alloca stmt to its offset within the per-thread slice.ad_stack_heap_base_llvm_— cached SSA value of the heap base pointer, emitted once inentry_block.init_offloaded_task_functionpre-scans the task body before any codegen runs and populates the first two, so that later sibling allocas never shift an earlier alloca's offset out from under a cached SSA pointer.visit(AdStackAllocaStmt)now emits:linear_thread_idxis the arch-appropriate invocation id (RuntimeContext::cpu_thread_idon CPU;block_idx * block_dim + thread_idxon CUDA / AMDGPU), matching howrand_statesis indexed.The old 256 KB function-scope budget guard (introduced in the previous PR) is deleted; its
ad_stack_fn_scope_bytes_accumulator is gone too. Heap-backed storage makes the ceiling irrelevant.OffloadedTaskgains anAdStackSizingInfo ad_stacksub-struct that propagates sizing to the host launcher:per_thread_stride,static_num_threads,dynamic_gpu_range_for, plus const values and gtmps byte offsets for range-forbegin/end.Per-arch codegen tweaks
codegen_cpu.cpp— fillscurrent_task->ad_stackwith the pre-scanned stride and setsstatic_num_threads = cpu_thread_id_range(CPU thread count is known at compile time).codegen_cuda.cpp— fillscurrent_task->ad_stack.static_num_threads = grid_dim * block_dimfor const-bound tasks, and marksdynamic_gpu_range_for = true+ recordsbegin_offset_bytes/end_offset_bytes/begin_const_value/end_const_valuefor dynamic range-for tasks so the launcher can resolve the actual iteration count at launch time.codegen_amdgpu.cpp— same as CUDA.Runtime (
llvm_runtime_executor.{h,cpp},runtime.cpp)LLVMRuntimegains two new fields:Ptr adstack_heap_buffer = nullptr; u64 adstack_heap_size = 0;. These are read by every adstack-backed task on the device side; the host writes to them through the cached field-pointer addresses.LlvmRuntimeExecutor::ensure_adstack_heap(needed_bytes):needed_bytes == 0 || needed_bytes <= adstack_heap_size_.new_size = max(needed_bytes, 2 * adstack_heap_size_)(amortised doubling).llvm_device()->allocate_memory), wraps in aDeviceAllocationGuard.runtime->adstack_heap_bufferandruntime->adstack_heap_size, caches them.memcpy_host_to_device(CUDA / AMDGPU) or plain pointer stores (CPU) against the cached addresses — no per-grow kernel launch.DeviceAllocationGuardvia move-assignment. Safety of the release (see the detailed block comment in the.cppand the matching field comment in the.h): CPU usesstd::free(trivially safe); CUDAcuMemFree_v2synchronises before returning; AMDGPUdealloc_memorypools throughCachingAllocator::releasewithout sync, and cross-launch safety on AMDGPU is provided by the synchronoushipFree(context_pointer)at the tail ofamdgpu::KernelLauncher::launch_llvm_kernel(the latent-fix in [AutoDiff] Autodiff 11: Latent adstack-adjacent fixes (AMDGPU hipFree, flush() keeps ctx_buffers_, always-preallocate) #536).get_runtime_temporaries_device_ptr()— cached lookup ofruntime->temporaries, used by the GPU launchers to read back dynamic range-for bounds.Per-arch launchers
runtime/cpu/kernel_launcher.{h,cpp}—Contextgains a parallelad_stack_needed_bytesvector, precomputed at register time (CPU sizing is static).launch_offloaded_taskscallsensure_adstack_heapper task.runtime/cuda/kernel_launcher.cpp— addsresolve_num_threads(task)which DtoH-memcpysbegin/endfromruntime->temporariesfor dynamic range-for tasks; callsensure_adstack_heapper task.runtime/amdgpu/kernel_launcher.cpp— same as CUDA.runtime/cuda/graph_manager.cpp— hard-errorsgraph=Trueon kernels where any task hasper_thread_stride > 0. Graph baking precludes host-side intervention between dispatches.Capacity knob (
compile_config.h,python/export_lang.cpp)default_ad_stack_sizeraised from 32 to 256.qd.init()kwarg. The comment block is rewritten to reflect the new heap-backing reality.ad_stack_size(per-stack explicit capacity) is unchanged.Docs (
docs/source/user_guide/autodiff.md)Drops the "SPIR-V on-chip cap" limitation from the known-limitations list (that was about the prior Function-scope SPIR-V path; with the SPIR-V heap landing in the next PR, it's gone too). Adds a "Tuning the capacity" section explaining
default_ad_stack_sizevsad_stack_sizeand the K+2-pushes-per-iteration rule for picking N.Tests (
tests/python/test_adstack.py)Heap-specific additions:
test_adstack_heap_grow_on_demand— two launches at increasing capacity pinpoint that the amortised-doubling grow path fires and the second launch reuses the bigger slab.test_adstack_heap_backed_exceeds_old_threadstack_budget— a kernel whose per-thread adstack bytes exceed the pre-PR 256 KB ceiling now compiles and runs correctly.test_adstack_cuda_graph_rejected_with_adstack—graph=Trueon an adstack kernel raises.Side-effect audit
AdStackSizingInfofields serialised via the existingOffloadedTaskkey hashStmtfields (all sizing lives in codegen /OffloadedTask)per_thread_stride > 0hipFree(context_pointer)tail.cppand.hcommentsper_thread_stride == 0path short-circuits the launcher-sideensure_adstack_heapStack
Split 2/3 of the former "heap-backed adstack" PR. Based on #536 (latent fixes). Followed by #493 (SPIR-V heap).