[SPIRV] Vulkan SPIR-V correctness: atomic-view aliasing, PSB stride, narrow storage caps, u1 cast, per-init layer recheck - #513
Conversation
|
Semi-orthogonal, but related: I wonder if we should start considering a CI that runs on a better GPU. If we want CI for Genesis, we will certainly need this. AMD GPU cloud does provide such GPUs I think. (or there is packet.ai, that @v01dXYZ discovered) |
|
Opus summary: SummaryFixes a SPIR-V codegen bug where plain loads/stores of float buffer elements could alias incorrectly with atomic float operations ( Root causeIn SPIR-V / Vulkan, each m.grad[i][j, k] += loss.grad
tmp = m.grad[i][j, k]
m.grad[i][j, k] = 0
n.grad += tmp * factorhits this exactly: the load reads the stale zero initial value, Behavior matrix
Good points
Bad points / risks
|
|
Most of hte bad points surfaced by Opus above look like AI-flation to me. But I'm wondering your thoughts on the following observation? : "Asymmetry between load and store for u1: load uses u8, store uses i8. This matches the previous behavior, but the helper does not encode it — store_buffer still |
cc6d8ea to
09976fc
Compare
ab96033 to
547a822
Compare
|
@claude review |
There was a problem hiding this comment.
My prior concern about cross-view aliasing on the CAS and non-add float atomic paths is resolved by fix #3 (the Aliased decoration on multi-view buffers in get_buffer_value) — that is the right fix. Given this now bundles 6 independent SPIR-V correctness fixes across codegen and the Vulkan device creator, it warrants a human look.
Extended reasoning...
Overview
The PR has grown from a single fix (route plain float loads/stores through the native float buffer view) to six bundled SPIR-V correctness fixes: (1) native float view for load/store; (2) ArrayStride decoration on PSB scalar/vector pointers; (3) Aliased decoration on every buffer that gets a second type view; (4) StorageBuffer{8,16}BitAccess capabilities plus SPV_KHR_{8,16}bit_storage extensions, gated on queried Vulkan feature bits; (5) route u1→i8 store widening through IRBuilder::cast (OpSelect) instead of spec-invalid OpBitcast %char %bool; (6) move validation-layer availability check above the cached-VkInstance short-circuit in create_instance. Touches spirv_codegen.cpp, spirv_ir_builder.cpp, the TaskCodegen header, rhi_constants.inc.h, and vulkan_device_creator.cpp.
Security risks
None that I can identify. These are correctness fixes in the SPIR-V code generator and the Vulkan device creator; they do not touch auth, crypto, permissions, or user-input boundaries. The narrow-type storage capabilities are only emitted when the underlying Vulkan device feature is queried true, so they cannot be forged into a shader that the driver would reject only at pipeline creation.
Level of scrutiny
High. SPIR-V codegen is production-critical — miscompiles here manifest as silent gradient corruption (fix #1 / #3), wrong ndarray reads (fix #2), pipeline-compiler crashes (fix #5), or spec-invalid shaders that validate only on lenient drivers (fix #4). The aliasing argument is subtle, and the decoration strategy in get_buffer_value relies on the invariant that every at_buffer call funnels through this path. No new tests are added — every fix leans on an existing regression. Worth a human with SPIR-V / Vulkan context walking through the aliasing argument and confirming the decoration site sweeps in every pairing (native-atomic vs CAS-emulation atomic vs plain load, per buffer, per type).
Other factors
My prior flagged bug has been directly addressed by fix #3, which is a meaningful progression. The bug hunting system found no new issues on this revision. The PR is labeled lowpri but the changes are not — they are correctness fixes with concrete reproducers documented in the description, which is exactly the kind of PR that benefits from a human signing off on the mechanism rather than shadow-approval.
|
checklist:
=> ok to merge |
4876859 to
731c547
Compare
…AccessChain scales correctly SPV_KHR_physical_storage_buffer requires an explicit `ArrayStride` decoration on `OpTypePointer PhysicalStorageBuffer` when the pointer is used with `OpPtrAccessChain`: the `Element` index is multiplied by that stride to produce the byte offset from the base address. Without the decoration the stride is undefined, and strict drivers collapse every indexed access back to the base - every `arr[i]` read returns `arr[0]` across the whole kernel, and every indexed ndarray write lands on slot 0. Narrow the fix to scalar/vector pointees. Struct and array pointees already carry explicit layout decorations (each member's `Offset`, array `ArrayStride`), so adding a top-level `ArrayStride` on the pointer to those is redundant; for scalars/vectors the natural stride is just the pointee's byte size. Pointers in Uniform / StorageBuffer / Input / Output / Workgroup storage classes don't use PSB arithmetic at all, so the decoration is skipped there. This stacks on the load/store aliasing fix: `pick_buffer_access_type` unblocks reverse-mode atomics on devices with `shaderBufferFloat32AtomicAdd`; the stride decoration unblocks indexed reads/writes through PSB on every Vulkan device. The two together drop our Vulkan test failure count from 506 to 84 across `tests/python/` (excluding the adstack / ndarray suites, which go from 12 failing to fully green on their own). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…ss-view accesses are ordered Review feedback on the native-float-view fix: that fix only closes the aliasing hazard when the atomic happens to use the native float view, which is true only for `OpAtomicFAddEXT` on devices that advertise `shaderBufferFloat32AtomicAdd` / `shaderBufferFloat64AtomicAdd` / `shaderBufferFloat16AtomicAdd`. Three other paths still emit the atomic through the u32-punned view while plain load/store go through the native float view: - CAS emulation on devices without the EXT (NVIDIA T4, our CI runner among them): `visit(AtomicOpStmt)` routes f32-add through `at_buffer(dest, get_quadrants_uint_type(dt))` and `atomic_operation` then runs `OpAtomicLoad` / `OpAtomicCompareExchange` on the u32 view. The plain load now on the float view is unordered against that CAS loop. - Non-add float atomics (min / max / mul) on every device: these never go through `OpAtomicFAddEXT`, always take the CAS path, always bind to the u32 view. - f16 / f64 add when `spirv_has_atomic_float16_add` / `spirv_has_atomic_float64_add` is not set: same CAS-on-u32 path. Close all of them structurally with an `Aliased` decoration on every buffer `OpVariable` that gets a second type view. `Aliased` is the SPIR-V signal that accesses through a variable may touch the same memory as accesses through another variable in the same storage class -- the driver must therefore preserve ordering across views, which is exactly what we need for the load-and-clear reverse-mode pattern to read back a freshly-atomic-added gradient. Decorate lazily: single-view buffers stay un-decorated so the compiler can still apply cross-variable scheduling on them. The decoration is applied when (and only when) a second distinct type view is minted, and it covers the newly-minted view plus every pre-existing peer in one sweep (tracked via `buffer_views_by_buffer_` + `aliased_decorated_buffer_ids_` to avoid emitting the decoration twice on the same id). `test_ad_dynamic_index.py::test_matrix_non_constant_index[arch=vulkan]` still passes; the wider Vulkan sweep on `tests/python/` (with `test_adstack.py` / `test_ndarray.py` excluded for independence) stays at 83 failing / 1778 passing, same delta as the native-float-view fix alone on this device. The decoration doesn't change the count on an AMD RX 7900 XTX because that device exposes `shaderBufferFloat32AtomicAdd` and hits the native-float-view path for every failing case in the suite; the decoration's job is to keep the CAS / non-add / f16-f64-no-native-add paths correct on devices that don't. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
… routing to an explicit float whitelist
Review feedback on the native-float-view fix:
- Switching `f16` loads / stores to the native `f16` descriptor view
requires the `CapabilityStorageBuffer16BitAccess` SPIR-V capability
(paired with the `SPV_KHR_16bit_storage` extension). That cap was
never emitted by `spirv_ir_builder.cpp`, so routing `f16` native
produced shaders that strict Vulkan drivers reject at pipeline
creation -- and even for the pre-existing uint-punning path, every
`i16` / `u16` load through a StorageBuffer already required the
same capability on spec-strict drivers, so the gap was latent
regardless. The same story applies to `i8` / `u8` and
`CapabilityStorageBuffer8BitAccess`.
- The `is_real(dt)` predicate in `pick_buffer_access_type` is too
open-ended: a future real-like primitive (bfloat16, an fp8 variant,
anything else `is_real` eventually admits) would silently fall into
the native-view branch before its storage-capability story has been
audited.
Fix both in one commit:
1. Add `spirv_has_storage_buffer_{8,16}bit_access` device caps to
`rhi_constants.inc.h`. Query them from the existing
`VkPhysicalDevice{8,16}BitStorageFeatures` structs in
`vulkan_device_creator.cpp`, strictly gated on
`storageBuffer{8,16}BitAccess` feature bits (the Vulkan 1.2-core
`VK_KHR_{8,16}bit_storage` promotion does not imply the feature is
supported, matching the pattern already used for
`bufferDeviceAddress`).
2. Emit `CapabilityStorageBuffer{8,16}BitAccess` + the matching
`SPV_KHR_{8,16}bit_storage` extension in the SPIR-V header
whenever the device caps are set. Unconditional relative to the
current kernel's type use -- the header cost is one extra
`OpCapability` + `OpExtension` per bit width, negligible against
the benefit of spec-compliant narrow StorageBuffer access for every
kernel that declares an `i8` / `i16` / `f16` field or ndarray.
3. Replace the `is_real(dt)` branch in `pick_buffer_access_type` with
an explicit `{f16, f32, f64}` whitelist. The three primitive
floats that exist today are the ones we've audited the
storage-capability story for; anything else must be added here
deliberately.
All existing Vulkan tests that exercised the uint-punned narrow
storage access (`test_ndarray.py` dtype parametrizations on `i8` /
`i16` / `f16`, etc.) retain their current behavior because the
capability emission is additive and gated on a feature the device
already exposes. The native-view routing scope shrinks from "every
real type" to "every real type we currently support", which
eliminates the silent-regression surface for future types.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
… accepts the cast chain
`TaskCodegen::store_buffer` widened a `u1` (bool) value to its backing
`i8` slot by emitting `OpBitcast %char %bool_val`. That's ill-formed
SPIR-V: `OpBitcast` requires a numerical scalar / vector or pointer
operand and rejects booleans (`OpTypeBool` has no defined bit pattern,
so a bitcast can't be defined either). `spirv-val` flagged the exact
message:
Expected input to be a pointer or int or float vector or scalar: Bitcast
%46 = OpBitcast %char %tmp3_u1
Most drivers just crash in the pipeline compiler on the ill-formed
input rather than surface a validation error. On Mesa RADV (AMD RX
7900 XTX) the symptom was a hard `SIGSEGV` deep inside
`libvulkan_radeon.so::create_compute_pipeline` the moment any kernel
storing to a `u1` field / ndarray / struct member was registered.
Python-side, that looked like `timeout: the monitored command dumped
core` on every one of the `test_tensor_consistency`,
`test_pickle[vulkan-u1]`, `test_matrixfree_{cg,bicgstab}`,
`test_struct_field_with_bool`, `test_dual_return_spirv`, and
`test_offload_cross` tests.
Route the `u1 -> i8` widening through `IRBuilder::cast`, which lowers
`bool -> int` to the canonical `OpSelect(cond, 1, 0)` with the 1 / 0
constants produced at the target integer type. That's the same route
`load_buffer`'s reverse path already uses on the read side, matches
what `IRBuilder::cast` does in every other `u1` context in the
codegen, and preserves the "`u1` serialises as 0 / 1" behaviour every
`to_numpy()` / `from_numpy()` user depends on. Non-`u1` stores keep
the existing `OpBitcast` path unchanged, so no other type path
changes.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…or init, not just the first `VulkanDeviceCreator::create_instance` normalises `params_.enable_validation_layer` against `check_validation_layer_support()` - flipping it to `false` when the host has no `VK_LAYER_KHRONOS_validation` - and then later gates the `spirv_has_non_semantic_info` cap and the logical-device layer array on that flag. The normalisation used to live below the cached-`VkInstance` short-circuit that the same function takes to work around an NVIDIA driver bug around repeated `vkDestroyInstance`/`vkCreateInstance` cycles (the instance is kept alive in the `VulkanLoader` singleton for process lifetime). Every re-init after the first one then read `params_.enable_validation_layer` as `true` (whatever the caller passed in, usually `config.debug`), jumped straight to the cached-instance return, and skipped the flip entirely - even on a host where the first `create_instance` had flipped it to `false`. The visible symptom in `test_overflow.py` / `test_print.py` was the first parametrisation passing (correct cap = 0) and every subsequent one in the same pytest session failing (stale cap = 1). Each test's `@test_utils.test(... debug=True)` wrapper calls `qd.reset()` + `qd.init(...)` which reconstructs the `VulkanDeviceCreator` with fresh params, so the first cycle flipped to `false`, but the second / third / ... cycles preserved `true`. The `spirv_has_non_semantic_info` cap then got set, the `NonSemantic.DebugPrintf` extinst was emitted in the shader, the validation layer wasn't actually loaded to intercept the output, and the `capfd`-based assertion against the `Addition overflow detected` string fails with an empty capture buffer. Move the layer-availability check above the cached-instance return. Re-running `vkEnumerateInstanceLayerProperties` on every cycle is microseconds and keeps the flag consistent whether the instance was freshly created or reused. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
731c547 to
9fa89c1
Compare
* [Misc] Warn user to disable caching when print_ir/QD_DUMP_IR enabled (Genesis-Embodied-AI#425) Co-authored-by: v01dxyz <v01dxyz@v01d.xyz> * [Build] Pin torch version to CUDA 12.8 for CUDA tests (Genesis-Embodied-AI#428) * [Misc] Fixing up taichi-dev urls (Genesis-Embodied-AI#429) * [Perf] Rename cuda_graph to gpu_graph across the codebase (Genesis-Embodied-AI#430) * Misc: fix typo integeral -> integral (Genesis-Embodied-AI#434) Co-authored-by: v01dxyz <v01dxyz@v01d.xyz> * [Perf] CUDA graph 4: call from multiple locations (Genesis-Embodied-AI#420) * [Bug] Fix fastcache not restoring graph_do_while_arg (Genesis-Embodied-AI#435) * [Perf] Cache last-call result in perf_dispatch for single-compatible case (Genesis-Embodied-AI#438) * Fix gpu_graph fallback on old Nvidia GPU. (Genesis-Embodied-AI#443) * Fix shared memory offset not reset between CUDA kernels. (Genesis-Embodied-AI#442) * [Misc] Allow disabling GPU graph via QD_GPU_GRAPH=0 env var (Genesis-Embodied-AI#439) * [Misc] Add named top-level loops (Genesis-Embodied-AI#440) * [Misc] Rename gpu_graph to graph (Genesis-Embodied-AI#446) * [Misc] Add cross-platform shuffle (Genesis-Embodied-AI#447) * [Bug] Fix graph_do_while on Windows: search for cudadevrt.lib (Genesis-Embodied-AI#456) * [Bug] Also search default CUDA toolkit install location on Windows (Genesis-Embodied-AI#461) * [SPIRV] Feature Parity Atomics & Shared Array (Genesis-Embodied-AI#432) * [Misc] Change clang format to 120 characters (Genesis-Embodied-AI#463) * [Misc] CUDA graph 5 Add fatbin (Genesis-Embodied-AI#464) * [Bug] Reuse VkInstance across init/reset cycles (Genesis-Embodied-AI#465) * [Perf] Tiles 1: _load, _store, _eye_ (Genesis-Embodied-AI#466) * [Misc] Remove dead InternalFuncStmt type_check override (Genesis-Embodied-AI#471) * [Perf] Tiles 2: add cholesky and ger (Genesis-Embodied-AI#472) * [Perf] Tiles 2b: add triangular solve (Genesis-Embodied-AI#474) * [Misc] Refactor: use _get_col/_set_col in tiles load/store/init (Genesis-Embodied-AI#475) * [Build] Fix flaky test_clock_accuracy (Genesis-Embodied-AI#436) * Fix AARCH64 emitting invalid asm in CUDA kernels. (Genesis-Embodied-AI#473) Co-authored-by: Hugh Perkins <hughperkins@gmail.com> * [AMDGPU] Enable HIP memory pool and surface pool-exhaustion errors. (Genesis-Embodied-AI#485) * [AMDGPU] Scope hsaco tmp dir per-user to avoid collisions. (Genesis-Embodied-AI#484) * [Perf] Tiles 3: Add slice syntax, qd.outer() and initial doc (Genesis-Embodied-AI#477) * [AMDGPU] Fix gradient computation. (Genesis-Embodied-AI#486) * Enable all backends that are supported in unit tests. (Genesis-Embodied-AI#488) * Fix SPIRV ID overflow for large kernels due to autodiff. (Genesis-Embodied-AI#489) * [Misc] Fix purity checker to allow accessing constants from quadrants modules (Genesis-Embodied-AI#487) * [Misc] Increase tolerance for clock monotonic test (Genesis-Embodied-AI#492) * [CI] Serialize api doc workflow (Genesis-Embodied-AI#494) * [CI] Increase tolerance for clock test (Genesis-Embodied-AI#506) * [CI] Increase clock test tolerance to 20% (Genesis-Embodied-AI#509) * [Perf] Add tensor_type parametrization to tile16 tests (Genesis-Embodied-AI#504) * [Perf] Tiles 4b: Migrate tiles16 tests to enable fastcache (Genesis-Embodied-AI#505) * [Perf] Tiles 4c: add Tiles16x16 proxy (Genesis-Embodied-AI#507) * [Perf] Tiles 4d: Consolidate slice error tests using parametrize (Genesis-Embodied-AI#508) * [Perf] Tiles 4: add SharedArray slice support (Genesis-Embodied-AI#482) * [Perf] Tiles 5: add Cholesky benchmark demo (Genesis-Embodied-AI#483) * [Doc] Add user guide page for subgroup shuffle (Genesis-Embodied-AI#512) * [Perf] Implement cross-platform shuffle_down (Genesis-Embodied-AI#510) * [Perf] Add portable subgroup reduce_add and reduce_all_add (Genesis-Embodied-AI#511) * [Perf] Add first warmup config to perf dispatch (Genesis-Embodied-AI#422) * [AutoDiff] Autodiff 1: Add baseline adstack regression test for unary_collections (Genesis-Embodied-AI#500) * [AutoDiff] Autodiff 2: Implement derivative for tan (Genesis-Embodied-AI#501) * [AutoDiff] Autodiff 3: Recompute tanh/exp on the operand in the reverse pass (Genesis-Embodied-AI#502) * [AutoDiff] Autodiff 4: Mark rsqrt as non-linear for adstack promotion (Genesis-Embodied-AI#503) * [AutoDiff] Autodiff 5: Fix adjoint-alloca placement for GlobalLoads outside the current range-for (Genesis-Embodied-AI#496) * [AutoDiff] Autodiff 6: Adstack regression tests (Genesis-Embodied-AI#491) * [AutoDiff] Autodiff 7: Fix header size in AdStackAllocaStmt to match u64 runtime layout (Genesis-Embodied-AI#534) * [AutoDiff] Autodiff 8: Surface LLVM adstack push/pop overflow as a Python exception (Genesis-Embodied-AI#535) * [AutoDiff] Autodiff 9: Guard against LLVM worker-thread stack overflow from large per-task adstack budget (Genesis-Embodied-AI#495) * [AutoDiff] Autodiff 10: Implement adstack for SPIR-V (Genesis-Embodied-AI#490) * [AutoDiff] Autodiff 11: Latent adstack-adjacent fixes (AMDGPU hipFree, flush() keeps ctx_buffers_, always-preallocate) (Genesis-Embodied-AI#536) * [Doc] Add AGENTS.md with instructions for AI agents (Genesis-Embodied-AI#541) * [Bug] Abort kernel execution on assertion failure instead of segfaulting (Genesis-Embodied-AI#419) * [Type] ndarray typing 1: Add eval_str=True to inspect.signature() calls (Genesis-Embodied-AI#411) * [CI] Suppress reportPrivateImportUsage in torch-using files (Genesis-Embodied-AI#552) * [Misc] QD_DUMP_IR dumps to files with the task_id added to the filename (Genesis-Embodied-AI#441) * [Type] ndarray typing 2: Fix NDArray single-arg subscript crash (Genesis-Embodied-AI#412) * [Test] Flush xdist channel before worker exit so test failure reports are visible (Genesis-Embodied-AI#555) * [CI] Reduce test retries on CI from 3 to 1. (Genesis-Embodied-AI#554) * [AutoDiff] Autodiff 12: Heap-backed adstack on LLVM backends (CPU/CUDA/AMDGPU) (Genesis-Embodied-AI#537) * [AutoDiff] Autodiff 13: Heap-backed adstack on SPIR-V backends (Metal, Vulkan) (Genesis-Embodied-AI#493) * [AutoDiff] Autodiff 14: Resolve bounded-inner-loop adstacks without default_ad_stack_size fallback (Genesis-Embodied-AI#539) * [SPIRV] Vulkan SPIR-V correctness: atomic-view aliasing, PSB stride, narrow storage caps, u1 cast, per-init layer recheck (Genesis-Embodied-AI#513) * [Build] Autodiff 15: Replace 2022 MoltenVK pin with LunarG Vulkan SDK fetch and sanitise MoltenVK cap advertisement (Genesis-Embodied-AI#551) * [Test] Suppress stock pytest-timeout to avoid conflict with pytest_hardtle (Genesis-Embodied-AI#557) * [Vulkan] Use SDK validation layer for debugPrintf instead of apt package (Genesis-Embodied-AI#562) * [Test] Fix flaky perf_dispatch tests by increasing work amounts (Genesis-Embodied-AI#559) * [Test] Add --maxfail CLI option to run_tests.py (default 20) (Genesis-Embodied-AI#558) * [CI] Vulkan debug printf fix to address flaky tests (Genesis-Embodied-AI#563) * [Docs] Add a new page to help for first time contributors (Genesis-Embodied-AI#426) Authored-by: v01dxyz <v01dxyz@v01d.xyz> * [AutoDiff] Autodiff 16: Resolve reverse-mode adstack depths per-launch via runtime-evaluated SizeExpr (Genesis-Embodied-AI#543) * Fix: raise error if device memory allocation fails (Genesis-Embodied-AI#451) (Genesis-Embodied-AI#453) Co-authored-by: v01dxyz <v01dxyz@v01d.xyz> Co-authored-by: Hugh Perkins <hughperkins@gmail.com> * [CI] Add CI job to check line wrapping of comments and docs (Genesis-Embodied-AI#564) * [Misc] Add coverage report to PRs, including kernels (Genesis-Embodied-AI#470) * [CI] CI wrap check feeds only diffs to agent (Genesis-Embodied-AI#567) * Skip 'flaky' test on MacOS CI. (Genesis-Embodied-AI#573) * [Test] Fix missing `import sys` in test_fail_device_memory_allocation (Genesis-Embodied-AI#574) * [CI] Fix Vulkan debugPrintf flake with session-scoped warmup (Genesis-Embodied-AI#571) * [AutoDiff] determine_ad_stack_size: replace whole-CFG Bellman-Ford with SCC + DAG DP (Genesis-Embodied-AI#575) * [Test] Fix macOS OOM skip reason to describe actual root cause (Genesis-Embodied-AI#576) * [Lang] whole_kernel_cse: 2.5x compile time speedup on large kernels (Genesis-Embodied-AI#577) * [CI] Add CI check for unnecessarily deleted comments (Genesis-Embodied-AI#570) * [CI] Migrate coverage report to github Check page (Genesis-Embodied-AI#566) * [Lang] Skip IR verifier between passes unless debug=true (Genesis-Embodied-AI#579) * [Lang] Inline AdStack ops on release LLVM codegen: dramatically reduces compile time for adstack-enabled reverse-mode kernels (Genesis-Embodied-AI#584) * [CUDA] Honor offline_cache=False end-to-end so QD_OFFLINE_CACHE=0 actually gives a cold compile (Genesis-Embodied-AI#580) * [Type] Tensor 24 (Genesis-Embodied-AI#561) Co-authored-by: hugh <hugh@slurm-login-0.slurm-login.tenant-slurm.svc.cluster.local> * [Lang] auto_diff host-walk reductions: dramatically faster front-end compile time on adstack-enabled reverse-mode kernels (Genesis-Embodied-AI#587) * [AutoDiff] Speed up reverse-mode kernel launches on GPU backends (Genesis-Embodied-AI#578) * [Vulkan] Move adstack-sizer scratch out of Function-scope memory to fix SPIR-V pipeline build failures (Genesis-Embodied-AI#588) * [AutoDiff] Improve diagnosis of unsupported reverse-mode AD patterns (Genesis-Embodied-AI#590) * [Bug] Fix: promote Ndarray to AnyArray in build_Name for flattened struct fields (Genesis-Embodied-AI#592) * [SPIR-V] Shrink reverse-grad kernel MSL by ~50% (Genesis-Embodied-AI#591) * [CI] Add CI check that PR changes have test coverage (Genesis-Embodied-AI#596) * [Perf] Enable zero-copy in to_torch() and to_numpy() (Genesis-Embodied-AI#450) * Add BufferView: safe sub-range ndarray access for kernels (Genesis-Embodied-AI#585) Co-authored-by: alanray-tech <alanray-tech@users.noreply.github.com> Co-authored-by: Hugh Perkins <hughperkins@gmail.com> * [Doc] Add user-facing fastcache documentation (Genesis-Embodied-AI#597) Co-authored-by: hugh <hugh@slurm-login-0.slurm-login.tenant-slurm.svc.cluster.local> * [Misc] Upgrade to enable v1 dlpack so to_numpy(copy=False) writable (Genesis-Embodied-AI#598) Co-authored-by: root <root@rtx-209-201.slurm-compute.tenant-slurm.svc.cluster.local> * [AutoDiff] Cut reverse-mode adstack memory usage 10x on all backends (Genesis-Embodied-AI#599) * [Misc] Add CI check for feature file factorization (Genesis-Embodied-AI#606) * [Perf] Skip _recursive_set_args for all-Field frozen dataclass structs (Genesis-Embodied-AI#607) Co-authored-by: Cursor <cursoragent@cursor.com> * [AutoDiff] SNode-arm bound-expr capture rejects fold-attack gate indices (Genesis-Embodied-AI#610) * [Misc] Suppress field fastcache warning for qd.Tensor (Genesis-Embodied-AI#615) Co-authored-by: Cursor <cursoragent@cursor.com> * [AutoDiff] Adstack heap: clip reducer count by per-task loop trip count (compile-time and SizeExpr-evaluated) (Genesis-Embodied-AI#611) * [Misc] Forward copy= through qd.Tensor, add copy=None option (Genesis-Embodied-AI#616) Co-authored-by: Cursor <cursoragent@cursor.com> * [Doc] Update README (Genesis-Embodied-AI#617) Co-authored-by: Cursor <cursoragent@cursor.com> * [CI] Fix coverage report showing def lines as uncovered (Genesis-Embodied-AI#623) Co-authored-by: Cursor <cursoragent@cursor.com> * [Perf] Generic launcher: persistent context, JIT-pointer reuse, Metal compute encoder, LLVM-GPU async memory ops (Part 1/2) (Genesis-Embodied-AI#619) * [CI] Encode Python-first testing policy in coverage-check prompt (Genesis-Embodied-AI#622) Co-authored-by: Cursor <cursoragent@cursor.com> * [CI] Add PR Line change report (Genesis-Embodied-AI#624) Co-authored-by: Cursor <cursoragent@cursor.com> * [CI] Disable quadrants pytest plugin during quadrants internal coverage runs (Genesis-Embodied-AI#629) Co-authored-by: Cursor <cursoragent@cursor.com> * [AutoDiff] Adstack load+store eliminations: EliminateRecomputableAdStackPushes pass + leaf extensions (Genesis-Embodied-AI#621) * [CI] Simplify coverage PR comment to a single linked line (Genesis-Embodied-AI#630) * [CUDA] Add AGX Thor, SM_110 (Genesis-Embodied-AI#631) Co-authored-by: Johnny Nunez and Hugh Perkins * [CI] Lines changed report: collapse PR comment to a single linked totals line (Genesis-Embodied-AI#632) * [FEATURE] Support external Metal command queue via qd.init (Genesis-Embodied-AI#618) Co-authored-by: Cursor <cursoragent@cursor.com> * [Perf] Cache adstack-sizer metadata per task across SPIR-V + LLVM-GPU; per-snode / DeviceAllocation invalidation (Part 2/2) (Genesis-Embodied-AI#620) * [AutoDiff] Disable EliminateRecomputableAdStackPushes pending mutated-SNode chain-leaf fix (Genesis-Embodied-AI#633) * [AutoDiff] Adstack chain-clone safety: mutated-SNode leaf reject + load_top consumer-aware guard (Genesis-Embodied-AI#634) * [Docs] Add user-guide page for qd.simt.block.* primitives (Genesis-Embodied-AI#638) * [Docs] Expand qd.simt.subgroup user-guide page to cover every op (Genesis-Embodied-AI#639) * [Perf] Streams 1-4 (Genesis-Embodied-AI#410) * [Docs] Add user-guide page for matrix decompositions and solvers (Genesis-Embodied-AI#643) * [Bug] Revert "[Perf] Streams 1-4 (Genesis-Embodied-AI#410)" (Genesis-Embodied-AI#650) * [Docs] Add user-guide page for atomics and bit operations (Genesis-Embodied-AI#640) * [Docs] Add user-guide page for qd.simt.grid.* primitives (Genesis-Embodied-AI#641) * [AutoDiff] Adstack max-reducer: parallel multi-axis MaxOverRange dispatch (Genesis-Embodied-AI#635) * [AMDGPU] Fix amdgpu parallel rand init (Genesis-Embodied-AI#658) * [Perf] Adstack: skip max-reducer recognizer on CPU + lift host-eval cap (Genesis-Embodied-AI#655) * [Perf] Re-land Streams 1-4 with bug fixes (Genesis-Embodied-AI#653) * [AMDGPU] Apply device_memory_GB=0.3 cap to AMDGPU tests (Genesis-Embodied-AI#659) * [Perf] Per-launch host sync: drop wait_idle on SPIR-V, pin stream and drop stream_synchronize on CUDA/AMDGPU (Genesis-Embodied-AI#654) * [AMDGPU] Unload hipModule_t in JITModuleAMDGPU destructor (Genesis-Embodied-AI#660) * [AMDGPU] Trim default mempool on qd.reset() (Genesis-Embodied-AI#669) * [AMDGPU] Hoist rand-state buffer to process lifetime (Genesis-Embodied-AI#668) * [Streams] Use events for streams serialization on AMDGPU and CUDA (Genesis-Embodied-AI#667) * [Perf] Adstack max-reducer: launch cache + zero-copy result map; content-stable registry_id (Genesis-Embodied-AI#671) * [SPIR-V] dispatch_max_reducers: register each task with the real kernel name (Genesis-Embodied-AI#675) * [AutoDiff] Debug-mode field/grad/dual: dtype, layout, and access-time invariants (Genesis-Embodied-AI#677) * [Docs] Add user-guide page for qd.algorithms.* device-wide algorithms (Genesis-Embodied-AI#642) Co-authored-by: alanray-tech <alan.ray@genesis-ai.company> * [Docs] Doc for existing atomics: switch support table to per-backend columns (Genesis-Embodied-AI#657) Co-authored-by: alanray-tech <alan.ray@genesis-ai.company> * [GPU] Cross gpu atomics (Genesis-Embodied-AI#666) Co-authored-by: alanray-tech <alan.ray@genesis-ai.company> * [GPU] Make block operations portable cross-gpu (Genesis-Embodied-AI#664) * [Perf] CPU LLVM adstack-cache: skip per-launch bump-writes + ndarray_shapes capture on forward-only handles (Genesis-Embodied-AI#685) * [GPU] Cross-GPU for grid ops (Genesis-Embodied-AI#670) * [Math] Make bitop operations portable cross-gpu (Genesis-Embodied-AI#662) * [AMDGPU] Always use wave64, on both RDNA and CDNA (Genesis-Embodied-AI#687) * [AMDGPU] Use syncscope("agent") for atomix xor to avoid CAS livelock (Genesis-Embodied-AI#672) * [GPU] New bit ops for QIPC (Genesis-Embodied-AI#679) * [GPU] Subgroup ops cross-gpu (Genesis-Embodied-AI#665) * [Graph] Rename CUDA Graph to Graph in docs (Genesis-Embodied-AI#691) * [SPIR-V] Fix FIFO-queue ordering when sharing command queue. (Genesis-Embodied-AI#694) * [Atomics] New QIPC ops for atomics (Genesis-Embodied-AI#690) * Pass dataclass sub-structs into qd.func (Genesis-Embodied-AI#698) * [AMDGPU] HIP graph runtime support for @qd.kernel(graph=True) (Genesis-Embodied-AI#692) * [CI] Add per-file timing report to Mac Metal test job (Genesis-Embodied-AI#695) Co-authored-by: Cursor <cursoragent@cursor.com> * [CI] Enable kernel disk cache during tests (Genesis-Embodied-AI#696) * [Math] New QIPC ops for single-threaded linalg (Genesis-Embodied-AI#683) * [BREAKING][GPU] New QIPC ops for subgroups (Genesis-Embodied-AI#676) * [GPU] New QIPC ops for block (Genesis-Embodied-AI#684) * [GPU] New device-level ops for QIPC (Genesis-Embodied-AI#693) * [algorithms] PrefixSumExecutor: drop unused GRID_SZ local (Genesis-Embodied-AI#701) * [block] sync(): fix unsupported-arch error message (Genesis-Embodied-AI#700) * [volatile_load] add qd.volatile_load primitive (closes Genesis-Embodied-AI#648) (Genesis-Embodied-AI#702) * [AutoDiff] Reject recycled identity_key in AdStackCache::register_adstack_sizing_info (Genesis-Embodied-AI#708) * [Vulkan] Declare GroupNonUniform SPIR-V caps and enable shaderSubgroupExtendedTypes (Genesis-Embodied-AI#707) * Fix duplicate HIP graph driver-function declarations after v1.0.0 merge The amd-integration fork had cherry-picked the HIP graph driver functions (graph_create / graph_destroy / graph_add_kernel_node / graph_instantiate / graph_exec_destroy / graph_launch), and upstream v1.0.0 added the same set. The per-file 3-way merge appended both copies into amdgpu_driver_functions.inc.h, producing redeclaration errors that broke the AMDGPU RHI/runtime compile. Drop the upstream duplicate block; the signatures are identical to the fork's existing declarations. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix AMDGPU launcher coherence and num_instructions visibility after v1.0.0 merge - kernel_launcher.cpp: the 3-way merge spliced upstream v1.0.0's launch_llvm_kernel rewrite (ephemeral arg/context buffers, explicit-stream path, AmdgpuDefaultStream PinGuard) onto the AMD fork's kernarg-by-value + persistent-scratch design, leaving references to undefined `ephemeral_context_ptr`. Restore the fork's coherent launch_llvm_kernel verbatim; it calls the (already merged) enhanced launch_offloaded_tasks, which keeps the max-reducer dispatch and stream-parallel groups adapted onto the AMD launch path. - llvm_context.h: both the fork and upstream added `num_instructions`; the merge kept upstream's private placement, but the AMDGPU codegen force-inline heuristic calls it statically from outside the class. Move it back to the public section. Co-authored-by: Cursor <cursoragent@cursor.com> * Restore async result D2H and hoist kernarg vectors in AMDGPU launcher The v1.0.0 merge resolution regressed two amd-integration baseline optimizations in launch_llvm_kernel / launch_offloaded_tasks: - The per-launch result-buffer copy was a blocking memcpy_device_to_host, forcing a host stall on every value-returning launch and serializing the GPU pipeline. Restore the async D2H (the caller synchronizes lazily when it needs the value); external-array transfers still stream_synchronize once before reading back. - launch_task constructed the kernarg std::vectors from initializer lists ({kernarg_payload} / {kernarg_size}) on every dispatch (heap alloc + free per launch). Hoist arg_ptrs/arg_sizes out of the per-task launch and reuse. Co-authored-by: Cursor <cursoragent@cursor.com> * amdgpu: default to LDS permlane64 emulation; drop host-x86 barrier asm on retarget Two AMDGPU JIT-compile crashes surfaced after the v1.0.0 merge pulled in the QIPC subgroup ops (Genesis-Embodied-AI#676), which made the rigid constraint solver's wave-cooperative reductions route through `amdgpu_cross_half_shuffle_i32`. Both manifested as a SIGSEGV inside `llvm::SIInstrInfo::getInstSizeInBytes` during `JITSessionAMDGPU::compile_module_to_hsaco` (i.e. at first kernel launch), and reproduce on gfx942 / MI300X. Baseline 0.4.6 never emitted these constructs, which is why it was unaffected. 1. Native `llvm.amdgcn.permlane64` lowering crashes the bundled LLVM 22.1.0 AMDGPU backend. Default `amdgpu_permlane64` to the existing LDS-roundtrip software emulation on every target (it produces identical results). Add `QD_AMDGPU_USE_NATIVE_PERMLANE64=1` to opt back into the native instruction once the backend bug is fixed; the old `QD_AMDGPU_FORCE_PERMLANE64_FALLBACK` is now the default and still honored. This is the actual crash fix. 2. The runtime module is compiled by the host x86_64 clang and only retargeted to amdgcn here, so `amdgpu_cross_half_shuffle_i32`'s `__asm__ volatile("" : "+v"(byte))` optimization barrier carries x86 flag clobbers (`~{dirflag},~{fpsr},~{flags}`) that are meaningless on AMDGPU. The IR verifies but the empty-body INLINEASM is invalid on the amdgcn target. Neutralize empty-body barrier asm during retarget (forward the tied value, then erase) so no stale host asm reaches codegen. On the wave64 targets we ship `ds_bpermute` already addresses the full wave, so the hint is a no-op. Co-authored-by: Cursor <cursoragent@cursor.com> * style: apply clang-format (v19.1.7) to AMDGPU fn_attrs and launcher sources CI pre-commit's clang-format hook reformatted these files (long declarations/lambda signatures collapsed onto single lines per the repo's clang-format config). Apply the same formatting so the hook passes. No functional changes. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(amdgpu): use CreateNeg for branchless i32 sgn instead of CreateSub(0, input) clang-tidy (modernize-use-nullptr, -warnings-as-errors) flagged `builder->CreateSub(0, input)` in the i32 sgn path: the literal `0` binds to the `llvm::Value*` LHS parameter as a null pointer, not an integer zero. Replace with `builder->CreateNeg(input)`, which emits `0 - input` with a proper zero constant -- identical intended semantics, and clang-tidy clean. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Robert Dazi <14996868+v01dXYZ@users.noreply.github.com> Co-authored-by: v01dxyz <v01dxyz@v01d.xyz> Co-authored-by: Hugh Perkins <hughperkins@gmail.com> Co-authored-by: Alexis DUBURCQ <alexis.duburcq@gmail.com> Co-authored-by: hugh <hugh@slurm-login-0.slurm-login.tenant-slurm.svc.cluster.local> Co-authored-by: alanray-tech <alan.ray@genesis-ai.company> Co-authored-by: alanray-tech <alanray-tech@users.noreply.github.com> Co-authored-by: root <root@rtx-209-201.slurm-compute.tenant-slurm.svc.cluster.local> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Johnny <johnnynuca14@gmail.com>
Vulkan SPIR-V correctness: atomic-view aliasing, PSB pointer stride, narrow-type storage caps,
u1storage cast, per-init validation-layer re-checkTL;DR
Six independent correctness gaps, all with observable silent-corruption failure modes on Vulkan:
1. Plain float loads / stores went through the
u32-punned view of the buffer instead of the nativef32viewOpAtomicFAddEXThas always used the nativef32view. PlainOpLoad/OpStorewent through theu32view. Each(buffer, element_type)pair gets its ownOpVariablewith a freshDescriptorSet/Binding, so those two views are different SPIR-V variables pointing at the sameVkBuffer. Without anAliaseddecoration, a driver is free to assume they don't alias, so the plain load is not memory-ordered against the preceding atomic. Reverse-mode AD's load-and-clear pattern (x.grad[i] += delta; tmp = x.grad[i]; x.grad[i] = 0; y.grad += tmp * factor) reads the stale zero, silent gradient drop.Originally discovered via
tests/python/test_ad_dynamic_index.py::test_matrix_non_constant_index[arch=vulkan]asserting0.0 == 1.0.2.
OpTypePointerinPhysicalStorageBufferwas missingArrayStrideOpPtrAccessChainon a PSB pointer multipliesElementbyArrayStrideto produce the byte offset. We were emitting bare pointers; strict drivers collapsed indexed access back to the base. Symptom: everyarr[i]ndarray read returnsarr[0].3. Type-punned views of the same buffer weren't flagged as mutually
AliasedFix 1 closes the aliasing hazard only for the native-atomic-add hot path. CAS-emulation atomic on devices without
shaderBufferFloat32AtomicAdd, non-add float atomics (min/max/mul), andf16/f64-without-native-add each route the atomic throughu32while plain load / store are now on the native float view -- same cross-view aliasing, different pairing. Close all of them with anAliaseddecoration on every bufferOpVariablethat gets a second type view. Lazy: single-view buffers stay undecorated.4. Narrow-type
StorageBufferaccess requiresStorageBuffer{8,16}BitAccesscapabilities that weren't emittedThe uint-punning path has always relied on
OpLoad/OpStorethroughu16/u8descriptor-bound pointers. PerSPV_KHR_16bit_storage/SPV_KHR_8bit_storage, that needsCapabilityStorageBuffer16BitAccess/CapabilityStorageBuffer8BitAccessin the header. Neither capability nor its extension was emitted, so strict drivers were within their rights to reject every Quadrants shader touching ani8/i16/f16/u8/u16field. Also narrow theis_real(dt)predicate inpick_buffer_access_typeto an explicit{f16, f32, f64}whitelist so a futurebfloat16/fp8doesn't silently fall through.5. Widening a
u1value to itsi8storage slot usedOpBitcast, which is ill-formed on booleansTaskCodegen::store_bufferwidenedu1throughOpBitcast %char %bool_val.spirv-val:Expected input to be a pointer or int or float vector or scalar: Bitcast. Mesa RADV (AMD RX 7900 XTX) crashed deep insidelibvulkan_radeon.so::create_compute_pipelinewith a rawSIGSEGVthe moment au1field / ndarray / struct member store was registered. Surfaced astimeout: the monitored command dumped coreacrosstest_tensor_consistency,test_pickle[vulkan-u1],test_matrixfree_{cg,bicgstab},test_struct_field_with_bool,test_dual_return_spirv, andtest_offload_cross.Route the
u1 -> i8widening throughIRBuilder::cast, which lowersbool -> intto the canonicalOpSelect(cond, 1, 0)at the target integer type. That's what the load side already does, and it preserves "u1serialises as 0 / 1" forto_numpy()/from_numpy().6.
VulkanDeviceCreator::create_instanceskipped the validation-layer re-check on re-initThe instance-reuse short-circuit (kept around an NVIDIA driver bug with repeated
vkDestroyInstance/vkCreateInstance) used to return beforecheck_validation_layer_support()ran, so every re-init after the first one readparams_.enable_validation_layeras whatever the caller passed (Truefordebug=True) instead of the flipped-Falsereflecting the host's actual layer availability. Downstream, thespirv_has_non_semantic_infocap followed the staleTrue, the shader emittedNonSemantic.DebugPrintfextinsts with no validation layer loaded to route them, and everycapfd-based assertion after the first parametrisation in the same pytest session failed with an empty capture buffer.test_overflow.py::test_shl_overflow[arch=vulkan-ty0-6]passing +test_shl_overflow[arch=vulkan-ty1-7]failing in the same-n 1session was the fingerprint.Move the check above the cached-instance return. Re-running
vkEnumerateInstanceLayerPropertieson every cycle costs microseconds and keeps the flag consistent.Why
Every fix is a spec-compliance gap, not a driver workaround:
shaderBufferFloat32AtomicAdd. CI's T4 falls back to CAS emulation, which matched the oldu32plain-load path, so the aliasing gap didn't exist there and the bug has been latent since native-atomic lowering landed on non-T4 drivers.u1store path. Dozens of existing tests were observed to segfault; the only reason this wasn't caught earlier is that Quadrants' otheru1code paths (loads, casts, atomics) already went throughIRBuilder::castand didn't exercise the broken store widening.debug=Truetest that expectsDebugPrintfoutput incapfdis at risk.None of the six regresses any other backend. Metal's MoltenVK path goes through SPIRV-Cross -> MSL, which ignores the aliasing question, the
ArrayStride, the narrow-storage caps, and the device-creator-path ordering. AMDGPU / CUDA / CPU don't use the SPIR-V codegen. Fix 5 is a bit-identical refactor on every other type path (onlyu1changes); fix 6 runs only insideVulkanDeviceCreator.Mechanism
Fix 1:
pick_buffer_access_typeroutes whitelisted float types through the native viewquadrants/codegen/spirv/spirv_codegen.cpp:Fix 2:
get_pointer_typeemitsArrayStridefor PSB scalar / vector pointeesquadrants/codegen/spirv/spirv_ir_builder.cpp::get_pointer_type. Whenstorage_class == PhysicalStorageBufferand the pointee is a primitive scalar / vector, decorate withArrayStride = sizeof(pointee). Struct / array pointees already carry full per-member layout, so no double-decoration. Non-PSB storage classes skipped.Fix 3:
get_buffer_valuedecorates multi-view buffer variables withAliasedquadrants/codegen/spirv/spirv_codegen.cpp::get_buffer_value. Track per-BufferInfolist of existing type views. When a second (or later) view is minted, retroactively decorate every peer withAliased; dedupe via an id-set. Single-view buffers stay undecorated.Fix 4: narrow-type storage caps gated on Vulkan-queried feature bits
Three-part change:
quadrants/inc/rhi_constants.inc.h: addspirv_has_storage_buffer_{8,16}bit_accessto theDeviceCapabilityenum.quadrants/rhi/vulkan/vulkan_device_creator.cpp: queryVkPhysicalDevice{8,16}BitStorageFeatures::storageBuffer{8,16}BitAccessand set the device caps. Strictly gated on the feature bit (the Vulkan 1.2-coreVK_KHR_{8,16}bit_storagepromotion doesn't imply the feature is supported).quadrants/codegen/spirv/spirv_ir_builder.cpp: emitCapabilityStorageBuffer{8,16}BitAccess+ the matchingSPV_KHR_{8,16}bit_storageextension when the device caps are set.Fix 5:
store_bufferroutesu1 -> i8throughIRBuilder::castquadrants/codegen/spirv/spirv_codegen.cpp::TaskCodegen::store_buffer. Widened the three-way logic:IRBuilder::cast(int, bool)already emitsOpSelect(cond, int_immediate(1), int_immediate(0))at the target type -- matches the spec-compliantbool -> intlowering, matches whatload_bufferdoes on the reverse path, and keeps the serialisation convention.Fix 6:
create_instanceruns the layer-availability check before the cached-instance returnquadrants/rhi/vulkan/vulkan_device_creator.cpp::create_instance. Moved:from after the cached-
VkInstanceshort-circuit to before it. Dropped the duplicate check that used to live lower in the function.Per-backend coverage matrix
u1cast)shaderBufferFloat32AtomicAddarr[i] -> arr[0]collapse on strict-PSB driversSPV_KHR_{8,16}bit_storagedriversu1store pipeline crash (RADVSIGSEGV; spec-invalidOpBitcast %bool)spirv_has_non_semantic_infoon re-init across a pytest sessionbool -> intcast was already correctTests
No new tests added in this PR. Every fix is pinned by an existing regression whose assertion already covered the failure mode:
tests/python/test_ad_dynamic_index.py::test_matrix_non_constant_index[arch=vulkan].tests/python/test_ndarray.py::test_ndarray_1d[arch=vulkan]+ every cross-suite test that doesarr[i]indexed access.test_ndarray.py's dtype-parametrised suite on narrow primitives.test_tensor_consistency,test_pickle[vulkan-u1],test_matrixfree_{cg,bicgstab},test_struct_field_with_bool,test_dual_return_spirv,test_offload_crossall disappear post-fix.test_overflow.py/test_print.pyfirst-passes-then-fails pattern across a single pytest session.Measured on an AMD RX 7900 XTX (advertises
shaderBufferFloat32AtomicAdd, strict PSB,storageBuffer{8,16}BitAccess), acrosstests/python/withtest_adstack.pyandtest_ndarray.pyexcluded for measurement independence: 506 failing pre-series, 22 failing post-series. Ontest_adstack.pyspecifically this takes Vulkan from 12 failing -> fully green; ontest_ndarray.py, from 3 failing -> fully green. The remaining 22 are all intest_scan.py-- a pre-series parallel-scan correctness issue unrelated to any of the six gaps here, left for a follow-up PR.Side-effect audit
pick_buffer_access_typereturnsget_quadrants_uint_type(dt)as beforeat_buffer(#1)ptr_val.stype.dt == u64branch preserved{f16, f32, f64}whitelist replacesis_real(dt)-- unknown reals fall into uint viewaliased_decorated_buffer_ids_id-setstorageBuffer{8,16}BitAccess(#4)u1stores (#5)OpBitcastpath as beforeu1loads (#5)IRBuilder::castpre-fixquadrants/codegen/spirv/,quadrants/rhi/vulkan/,quadrants/inc/untouched