Skip to content

Fix AARCH64 emitting invalid asm in CUDA kernels.#473

Merged
hughperkins merged 4 commits into
mainfrom
duburcqa/fix_aarch64_invalid_asm
Apr 14, 2026
Merged

Fix AARCH64 emitting invalid asm in CUDA kernels.#473
hughperkins merged 4 commits into
mainfrom
duburcqa/fix_aarch64_invalid_asm

Conversation

@duburcqa

@duburcqa duburcqa commented Apr 13, 2026

Copy link
Copy Markdown
Contributor

Problem

On AArch64 + CUDA (e.g. NVIDIA GB10 / Jetson Thor), any Genesis kernel that pulls in cuda_active_mask, cuda_match_all_sync_i32, or cuda_match_any_sync_i64 crashes during LLVM PTX code generation with:

error: couldn't allocate output register for constraint 'w' at line 37014

Root cause: runtime.cpp contains PTX inline assembly for these three functions, but uses #if aarch64 guards to select AArch64 register constraints ("=w" for 32-bit FP/SIMD registers) instead of x86 ones ("=r"). When clang compiles runtime.cpp
to bitcode on AArch64, it embeds the "=w" constraint. Later, when this bitcode is linked into an NVPTX module, the LLVM NVPTX backend rejects "=w" — it is a valid AArch64 constraint but meaningless in PTX.

Separately, the CPU backend also fails on AArch64 with Function "runtime_initialize" not found, because clang's default -moutline-atomics emits calls to __aarch64_ldadd4_acq_rel etc. that the ORC JIT linker cannot resolve.

Solution

Commit 1: Replace the three inline-asm functions with trivial stubs (return 0) and add corresponding patch_intrinsic calls in llvm_context.cpp to replace them with proper LLVM NVPTX intrinsics (nvvm_activemask, nvvm_match_all_sync_i32p,
nvvm_match_any_sync_i32) at module-init time. This is the same mechanism already used for dozens of other CUDA intrinsics in the same file. The old patch_intrinsic calls for these were commented out due to LLVM 10 issues that are long fixed in
LLVM 22.

Commit 2: Add -mno-outline-atomics to the clang flags that compile runtime.cpp to bitcode, so atomic operations use inline instructions instead of external helper calls.

Platform impact

  • AArch64 + CUDA: fixes the PTX codegen crash (the original bug).
  • AArch64 + CPU: fixes the JIT linker crash (outline atomics).
  • x86_64 + CUDA: no behavioral change. patch_intrinsic produces the same NVPTX IR that the old inline asm did. CI tests (test_active_mask, test_match_any, test_match_all in test_simt.py) cover this.
  • x86_64 + CPU: no change. -mno-outline-atomics is a no-op on x86.

Note on cuda_match_any_sync_i64

This function has a stub in runtime.cpp but is not registered in internal_ops.inc.h or type_system.cpp, so no codegen path emits calls to it. This was true before this PR as well (the old inline asm was also dead code). No patch_intrinsic is
added for it. A comment documents why.

@hughperkins

Copy link
Copy Markdown
Collaborator

Test process? I'm guessing this makes Genesis kernels pass, which used to fail? Could we add some relevant quadrants tests please.

@hughperkins

Copy link
Copy Markdown
Collaborator

Oh I guess we would need runners that are both:

  • arm
  • have a cuda gpu
    🤔

@hughperkins

Copy link
Copy Markdown
Collaborator

Could you provide a description please of:

  • what challenge you encountered?
  • on which specific platforms (is it just aarch + cuda?)
  • how this PR intends to solve this?
  • how we can be fairly sure that this PR won't break things on other platform combinations?

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — cleanly replaces architecture-guarded inline asm with LLVM intrinsics via the established patch_intrinsic pattern.

Extended reasoning...

Overview

Two files changed: llvm_context.cpp (adds three new patch_intrinsic calls for cuda_active_mask, cuda_match_all_sync_i32, and cuda_match_any_sync_i64) and runtime_module/runtime.cpp (removes the corresponding architecture-conditional asm volatile blocks, leaving plain stubs). The net effect is that AARCH64-host builds no longer emit PTX-incompatible inline assembly for these three CUDA warp-match / active-mask intrinsics.

Security Risks

None. This is internal GPU runtime JIT compilation code with no auth, crypto, or user-facing attack surface.

Level of Scrutiny

Medium. The code touches CUDA kernel compilation, where a wrong intrinsic or wrong argument order could silently produce bad PTX. That said, the implementation follows the exact pattern used for every other CUDA intrinsic in the same file. The cuda_match_all_sync_i32 case correctly handles the {i32, i1} struct return of nvvm_match_all_sync_i32p by extracting element 0, which is the only non-trivial piece of logic in the diff.

Other Factors

The one flagged bug (duplicate patch_intrinsic("block_memfence",...) call) is pre-existing and carries zero semantic impact — both calls produce identical IR. It was not introduced by this PR. No existing reviewer comments are outstanding. The change is small, self-contained, and directly addresses the stated AARCH64 asm-constraint mismatch.

Comment thread quadrants/runtime/llvm/llvm_context.cpp
@hughperkins

Copy link
Copy Markdown
Collaborator

my converastion with opus:

why didnt we use patch intrinscis before?

Opus:

  The old code on main tells the story clearly. There were two distinct reasons:


  1. `cuda_match_all_sync_i32` — return type mismatch

  The comments say "LLVM 10.0.0 seems to have a bug on this intrinsic function." Looking more closely, it wasn't really a bug — it's a type mismatch. The LLVM intrinsic
  nvvm_match_all_sync_i32p returns {i32, i1} (a struct with the matched value and a predicate), but the C stub cuda_match_all_sync_i32 is declared as returning plain u32. The
  simple patch_intrinsic lambda just does CreateRet(CreateIntrinsic(...)), which doesn't work when the intrinsic's return type doesn't match the function's return type. You'd
  need to ExtractValue first — which is exactly what this PR adds.
  So back in the LLVM 10 era, someone tried it, it didn't work, assumed it was an LLVM bug, and fell back to inline PTX assembly instead. The commented-out code also has a typo
  (nvvm_math_all_sync_i32 instead of nvvm_match_all_sync_i32p), suggesting it was never fully working.


  2. `cuda_match_any_sync_i64` and `cuda_active_mask` — likely same era workaround

  These were also commented out with the same "LLVM 10.0.0 bug" note. They may have genuinely been buggy in LLVM 10, or they may have been collateral — disabled at the same time
   as match_all and replaced with inline asm as a batch. The inline asm worked fine on x86_64, so nobody revisited it until the project needed to build on AARCH64, where the
  register constraints ("=w" vs "=r") are different and the old asm was wrong.


  TL;DR

  The original patch_intrinsic approach was tried but abandoned because (a) match_all returns a struct that the generic helper couldn't handle, (b) the other two were assumed
  buggy in the old LLVM version, and (c) inline asm worked fine on x86_64 so there was no pressure to fix it. The AARCH64 port exposed the inline asm as broken, and this PR
  finally does what should have been done all along — use the proper LLVM intrinsics, with a custom IR wrapper for the match_all struct return.

what does match all do?

Opus:

  cuda_match_all_sync_i32 is a CUDA warp-level primitive. Given a warp mask and a value, it checks whether all active threads (identified by the mask) hold the same value. It
  returns:
  • The matching value (as i32) — which is just the common value if all threads agree.
  • A predicate (i1) — true if all active threads had the same value, false otherwise.

  The LLVM intrinsic (nvvm_match_all_sync_i32p) returns both as a struct {i32, i1}, but this codebase only cares about the i32 part, which is why the PR adds the
  ExtractValue(result, {0}) wrapper to discard the predicate.
  Its sibling match_any_sync is the more general variant — it returns a mask of which threads share the same value as the calling thread, without requiring unanimity.
  These are used for things like warp-level reductions, deduplication, and divergence detection in GPU kernels.

is there anything that could go wrong in this pr, or as a rsul of merging this pr?

Opus:

1. `cuda_match_all_sync_i32` — the predicate is silently dropped

  The NVVM intrinsic nvvm_match_all_sync_i32p returns {i32, i1} where the i1 is a predicate indicating whether all threads actually matched. This PR extracts only element {0}
  (the i32 value) and discards the predicate. The Python API warp.match_all() also only returns one value.
  This means there's no way for user code to know whether all threads actually agreed. The returned i32 is the value itself, but without the predicate it's ambiguous — if
  threads didn't all match, the return value is unspecified per CUDA semantics. If anyone is relying on this function to detect unanimity (which is the whole point of match_all
  vs match_any), they'd have no way to do so. This isn't a regression though — the old inline asm also dropped the predicate.


  2. `cuda_match_any_sync_i64` — not registered as an internal op

  This is the most concrete concern. The function cuda_match_any_sync_i64 has a C stub in runtime.cpp and is now patched with patch_intrinsic in llvm_context.cpp, but it is not
  listed in internal_ops.inc.h or type_system.cpp. There's no CUDA_MATCH_SYNC(any, i64) and no PER_INTERNAL_OP(cuda_match_any_sync_i64). There's also no Python-level match_any
  variant that takes i64.
  This means the patch_intrinsic call for cuda_match_any_sync_i64 will likely never find a function in the module (since no codegen path emits calls to it), and the code is
  dead. Not harmful, but also not useful — and if someone later adds the i64 variant to the IR, they might not realize the C stub is a no-op return 0 and would silently get
  wrong results on non-CUDA backends.


  3. CPU backend — stubs return 0

  The stubs now unconditionally return 0 on all architectures. Before, on CUDA they used inline asm, and on non-CUDA they returned 0. The new code relies entirely on
  patch_intrinsic to replace the body at LLVM IR level for CUDA. If the patching somehow fails (e.g., function name mismatch, module doesn't contain the function), the stub
  returning 0 would silently execute on GPU with wrong results. The old inline asm would at least produce the correct PTX. That said, this failure mode is unlikely since
  match_any_sync_i32 already uses patch_intrinsic on main and works fine.


  4. Behavioral change for `cuda_active_mask` and `cuda_match_any_sync_i64`

  These were previously implemented with inline asm on CUDA and were actually functional at runtime. Now they rely on patch_intrinsic. If the patching works (and it should — the
   same pattern is used for dozens of other intrinsics), this is a clean improvement. But it is a change in code generation path, so any subtle differences in how LLVM lowers
  the intrinsic vs the inline asm could surface. For example, the inline asm was marked volatile, preventing reordering — the LLVM intrinsics may have different scheduling
  behavior.


  Summary

  • Real risk: Essentially none for correctness, assuming your LLVM version supports these intrinsics (you're on LLVM 22, so yes).
  • Dead code: The cuda_match_any_sync_i64 patch is unreachable since there's no IR/type registration for it.
  • Design gap (pre-existing): match_all discards the predicate, making it functionally questionable, but that's not a regression from this PR.

  The PR looks safe to merge. The main actionable feedback would be to either remove the dead cuda_match_any_sync_i64 code or properly register it in internal_ops.inc.h and
  type_system.cpp.

@hughperkins

Copy link
Copy Markdown
Collaborator

what tests are there for this function? how do we know it continues to work ok on cuda? Sorry on cuda x86

Opus:

  Test coverage

  There are three relevant tests in tests/python/test_simt.py, all gated to arch=qd.cuda:
  • `test_match_any` — verifies match_any returns correct thread bitmasks for two groups of values
  • `test_match_all` — verifies match_all returns 0xFFFFFFFF when all threads agree, and 0 when they don't
  • `test_active_mask` — verifies active_mask returns the correct mask for 16 active threads

  Plus a lower-level test_active_mask in test_cuda_internals.py.


  CI coverage on CUDA x86

  The test_gpu.yml workflow has a test_linux_cuda job that runs on gpu-t4-4-core runners (NVIDIA T4, x86_64). It runs:

  python tests/run_tests.py -r 1 -v --arch cuda

  This runs all Python tests under tests/python/ filtered to CUDA arch, which includes test_simt.py. The T4 is compute capability 75, so the >= 70 guards in test_match_any and
  test_match_all will pass and the tests will execute.


  So: yes, CI should catch regressions on CUDA x86

  As long as the GPU test workflow runs on this PR (it's triggered via workflow_call, so it depends on the parent workflow), those three tests will exercise all three changed
  functions on a real T4 GPU on x86_64. If the intrinsic patching produces different results than the old inline asm, the assertions would fail.
  The main gap is that there are no tests for `cuda_match_any_sync_i64` — but as I noted earlier, that function isn't wired up in the IR/type system anyway, so it's dead code on
   both main and this branch.

@hughperkins

Copy link
Copy Markdown
Collaborator

@duburcqa please could you address Opus's comments on cuda_match_any_sync_i64

@duburcqa

Copy link
Copy Markdown
Contributor Author

@hughperkins Addressed — cuda_match_any_sync_i64 is not registered in internal_ops.inc.h or type_system.cpp, so no codegen path emits calls to it. This was already the case before this PR (the old inline asm was dead code too). I've removed the patch_intrinsic call for it and added a comment explaining why it's skipped. The stub in runtime.cpp remains (harmless, and removing it would be a separate cleanup).

Updated the PR description with the full context: what broke, on which platforms, how it's fixed, and why it's safe on x86_64.

Force-pushed to update the commits.

@duburcqa
duburcqa force-pushed the duburcqa/fix_aarch64_invalid_asm branch from 40142f6 to e2e43db Compare April 13, 2026 10:48
@duburcqa

Copy link
Copy Markdown
Contributor Author

Test process? I'm guessing this makes Genesis kernels pass, which used to fail? Could we add some relevant quadrants tests please.

pytest -xsv -n 0 "tests/python/test_simt.py::test_active_mask" "tests/python/test_simt.py::test_match_any" "tests/python/test_simt.py::test_match_all"
[Quadrants] version 0.0.0, llvm 22.1.0, commit 13f18a5b, linux, python 3.12.3
================================================================================================================= test session starts ==================================================================================================================
platform linux -- Python 3.12.3, pytest-9.0.2, pluggy-1.6.0 -- /home/genesis-ai/workspace/src/genesis/.venv/bin/python3
cachedir: .pytest_cache
rootdir: /home/genesis-ai/workspace/src/quadrants/tests
configfile: pytest.ini
plugins: xdist-3.8.0, repeat-0.9.4, forked-1.6.0, rerunfailures-16.1, anyio-4.12.1, print-1.2.2, timeout-2.4.0, syrupy-5.1.0
collected 3 items                                                                                                                                                                                                                                      

tests/python/test_simt.py::test_active_mask[arch=cuda] [Quadrants] Starting on arch=cuda
error: couldn't allocate output register for constraint 'w' at line 37014

@duburcqa

Copy link
Copy Markdown
Contributor Author

For the other fix, any atomic unit test in tests/python/test_atomic.py could trigger the crash:

tests/python/test_atomic.py::test_atomic_add_global_i32[arch=arm64-1] [Quadrants] Starting on arch=arm64
JIT session error: Symbols not found: [ __aarch64_ldadd4_acq_rel, __aarch64_ldadd8_acq_rel, __aarch64_swp4_acq_rel, __aarch64_swp8_acq_rel ]
[E 04/13/26 13:11:42.189 211990] [jit_cpu.cpp:lookup_in_module@198] Function "runtime_initialize" not found


ERROR

======================================================================================================================== ERRORS ========================================================================================================================
______________________________________________________________________________________________ ERROR at setup of test_atomic_add_global_i32[arch=arm64-1] ______________________________________________________________________________________________

request = <SubRequest 'wanted_arch' for <Function test_atomic_add_global_i32[arch=arm64-1]>>, req_arch = <Arch.arm64: 1>, req_options = {'print_full_traceback': True}

    @pytest.fixture(autouse=True)
    def wanted_arch(request, req_arch, req_options):
        if req_arch is not None:
            if req_arch == qd.cuda:
                if not request.node.get_closest_marker("run_in_serial"):
                    # Optimization only apply to non-serial tests, since serial tests
                    # are picked out exactly because of extensive resource consumption.
                    # Separation of serial/non-serial tests is done by the test runner
                    # through `-m run_in_serial` / `-m not run_in_serial`.
                    req_options = {
                        "device_memory_GB": 0.3,
                        "cuda_stack_limit": 1024,
                        **req_options,
                    }
                else:
                    # Serial tests run without aggressive resource optimization
                    req_options = {"device_memory_GB": 1, **req_options}
            if "print_full_traceback" not in req_options:
                req_options["print_full_traceback"] = True
>           qd.init(arch=req_arch, enable_fallback=False, **req_options)

tests/python/conftest.py:73: 
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ 

arch = <Arch.arm64: 1>, default_fp = None, default_ip = None, _test_mode = False, enable_fallback = False, require_version = None, print_non_pure = False, src_ll_cache = True, kwargs = {}, current_dir = '/home/genesis-ai/workspace/src/quadrants'
cfg = <quadrants._lib.core.quadrants_python.CompileConfig object at 0xfb8d5fcbb830>, spec_cfg = <quadrants.lang.misc._SpecialConfig object at 0xfb8e19d37a70>, env_comp = <quadrants.lang.misc._EnvironmentConfigurator object at 0xfb8d5f2f9be0>
env_spec = <quadrants.lang.misc._EnvironmentConfigurator object at 0xfb8d5f2fa9f0>, env_default_fp = None, env_default_ip = None

    def init(
        arch=None,
        default_fp=None,
        default_ip=None,
        _test_mode: bool = False,
        enable_fallback: bool = True,
        require_version: str | None = None,
        print_non_pure: bool = False,
        src_ll_cache: bool = True,
        **kwargs,
    ):
        """Initializes the Quadrants runtime.
    
        This should always be the entry point of your Quadrants program. Most
        importantly, it sets the backend used throughout the program.
    
        Args:
            arch: Backend to use. This is usually :const:`~quadrants.lang.cpu` or :const:`~quadrants.lang.gpu`.
            default_fp (Optional[type]): Default floating-point type.
            default_ip (Optional[type]): Default integral type.
            require_version: A version string.
            print_non_pure: Print the names of kernels, at the time they are executed, which are not annotated with
                            @qd.pure
            src_ll_cache: enable SRC-LL-CACHE, which will accelerate loading from cache, across all architectures,
                          for pure kernels (i.e. kernels declared as @qd.pure)
            **kwargs: Quadrants provides highly customizable compilation through
                ``kwargs``, which allows for fine grained control of Quadrants compiler
                behavior. Below we list some of the most frequently used ones. For a
                complete list, please check out
                https://github.com/Genesis-Embodied-AI/quadrants/blob/master/quadrants/program/compile_config.h.
    
                * ``cpu_max_num_threads`` (int): Sets the number of threads used by the CPU thread pool.
                * ``debug`` (bool): Enables the debug mode, under which Quadrants does a few more things like boundary checks.
                * ``print_ir`` (bool): Prints the CHI IR of the Quadrants kernels.
                *``offline_cache`` (bool): Enables offline cache of the compiled kernels. Default to True. When this is enabled Quadrants will cache compiled kernel on your local disk to accelerate future calls.
                *``random_seed`` (int): Sets the seed of the random generator. The default is 0.
                *``debug_dump_path`` (str): used as the base path for QD_DUMP_IR and similar
        """
        # FIXME(https://github.com/taichi-dev/taichi/issues/4811): save the current working directory since it may be
        # changed by the Vulkan backend initialization on OS X.
        current_dir = os.getcwd()
    
        # Check if installed version meets the requirements.
        if require_version is not None:
            check_require_version(require_version)
    
        if "default_up" in kwargs:
            raise KeyError("'default_up' is always the unsigned type of 'default_ip'. Please set 'default_ip' instead.")
        # Make a deepcopy in case these args reference to items from qd.cfg, which are
        # actually references. If no copy is made and the args are indeed references,
        # qd.reset() could override the args to their default values.
        default_fp = _deepcopy(default_fp)
        default_ip = _deepcopy(default_ip)
        kwargs = _deepcopy(kwargs)
        reset()
    
        cfg = impl.default_cfg()
        cfg.offline_cache = True  # Enable offline cache in frontend instead of C++ side
    
        spec_cfg = _SpecialConfig()
        env_comp = _EnvironmentConfigurator(kwargs, cfg)
        env_spec = _EnvironmentConfigurator(kwargs, spec_cfg)
    
        # configure default_fp/ip:
        # TODO: move these stuff to _SpecialConfig too:
        env_default_fp = os.environ.get("QD_DEFAULT_FP")
        if env_default_fp:
            if default_fp is not None:
                _qd_core.warn(
                    f'Environment variable QD_DEFAULT_FP={env_default_fp} overridden by qd.init argument "default_fp"'
                )
            elif env_default_fp == "32":
                default_fp = f32
            elif env_default_fp == "64":
                default_fp = f64
            elif env_default_fp is not None:
                raise ValueError(f"Invalid QD_DEFAULT_FP={env_default_fp}, should be 32 or 64")
    
        env_default_ip = os.environ.get("QD_DEFAULT_IP")
        if env_default_ip:
            if default_ip is not None:
                _qd_core.warn(
                    f'Environment variable QD_DEFAULT_IP={env_default_ip} overridden by qd.init argument "default_ip"'
                )
            elif env_default_ip == "32":
                default_ip = i32
            elif env_default_ip == "64":
                default_ip = i64
            elif env_default_ip is not None:
                raise ValueError(f"Invalid QD_DEFAULT_IP={env_default_ip}, should be 32 or 64")
    
        if default_fp is not None:
            impl.get_runtime().set_default_fp(default_fp)
        if default_ip is not None:
            impl.get_runtime().set_default_ip(default_ip)
    
        # submodule configurations (spec_cfg):
        env_spec.add("log_level", str)
        env_spec.add("gdb_trigger")
        env_spec.add("short_circuit_operators")
        env_spec.add("print_full_traceback")
        env_spec.add("unrolling_limit")
    
        # compiler configurations (qd.cfg):
        for key in dir(cfg):
            if key in ["arch", "default_fp", "default_ip"]:
                continue
            _cast = type(getattr(cfg, key))
            if _cast is bool:
                _cast = None
            env_comp.add(key, _cast)
    
        unexpected_keys = kwargs.keys()
    
        if len(unexpected_keys):
            raise KeyError(f'Unrecognized keyword argument(s) for qd.init: {", ".join(unexpected_keys)}')
    
        if (cfg.print_ir or os.getenv("QD_DUMP_IR") == "1") and cfg.offline_cache:
            util.warning(
                "Even with print_ir/QD_DUMP_IR enabled, already cached kernels won't get their IRs shown. "
                "You might want to disable caching with offline_cache=False. "
                "[warning_code=DUMP_IR_CACHE_MISMATCH]"
            )
    
        # dispatch configurations that are not in qd.cfg:
        runtime = impl.get_runtime()
        if not _test_mode:
            _qd_core.set_core_trigger_gdb_when_crash(spec_cfg.gdb_trigger)
            runtime.short_circuit_operators = spec_cfg.short_circuit_operators
            runtime.print_full_traceback = spec_cfg.print_full_traceback
            runtime.unrolling_limit = spec_cfg.unrolling_limit
            runtime.src_ll_cache = src_ll_cache
            runtime.print_non_pure = print_non_pure
            _logging.set_logging_level(spec_cfg.log_level.lower())
    
        # select arch (backend):
        env_arch = os.environ.get("QD_ARCH")
        if env_arch is not None:
            _logging.info(f"Following QD_ARCH setting up for arch={env_arch}")
            arch = _qd_core.arch_from_name(env_arch)
        cfg.arch = adaptive_arch_select(arch, enable_fallback)
        print(f"[Quadrants] Starting on arch={_qd_core.arch_name(cfg.arch)}")
    
        if cfg.arch == _qd_core.amdgpu and get_os_name() == "win":
            _logging.warn("AMDGPU support on Windows is experimental and may not work as expected.")
    
        if _test_mode:
            return spec_cfg
    
        get_default_kernel_profiler().set_kernel_profiler_mode(cfg.kernel_profiler)
    
        impl.get_runtime()._arch = cfg.arch
    
        # create a new program (skip for python backend — no C++ runtime needed):
        if cfg.arch != _qd_core.python:
            impl.get_runtime().create_program()
            _logging.trace("Materializing runtime...")
>           impl.get_runtime().prog.materialize_runtime()
E           RuntimeError: [jit_cpu.cpp:lookup_in_module@198] Function "runtime_initialize" not found

python/quadrants/lang/misc.py:485: RuntimeError
=============================================================================================================== short test summary info ================================================================================================================
ERROR tests/python/test_atomic.py::test_atomic_add_global_i32[arch=arm64-1] - RuntimeError: [jit_cpu.cpp:lookup_in_module@198] Function "runtime_initialize" not found
!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! stopping after 1 failures !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
============================================================================================================== 1 passed, 1 error in 0.59s ==============================================================================================================
Bus error (core dumped)

@hughperkins

Copy link
Copy Markdown
Collaborator

Thanks! Seems there are some faliures in test_pointer https://github.com/Genesis-Embodied-AI/quadrants/actions/runs/24339347087/job/71067191548?pr=473 Re-running, in case it's an infra issue.

@duburcqa
duburcqa force-pushed the duburcqa/fix_aarch64_invalid_asm branch 2 times, most recently from 9d2001f to 9ddc9f2 Compare April 13, 2026 12:54
@duburcqa

duburcqa commented Apr 13, 2026

Copy link
Copy Markdown
Contributor Author

Fixed. Removing 'cuda_match_any_sync_i64' was a mistake. I documented this in the code.

@duburcqa

Copy link
Copy Markdown
Contributor Author

@claude review

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — the inline-asm → LLVM-intrinsic replacement is the right approach, the custom IR wrapper for cuda_match_all_sync_i32's struct return correctly extracts element 0, and -mno-outline-atomics cleanly fixes the CPU JIT crash.

Extended reasoning...

Overview

Three files changed: runtime.cpp removes arch-specific PTX inline asm and replaces three functions with trivial stubs; llvm_context.cpp adds corresponding patch_intrinsic calls (plus a custom IR builder for the struct-returning nvvm_match_all_sync_i32p); CMakeLists.txt adds -mno-outline-atomics to the bitcode compilation flags.

Security risks

None. This is a LLVM bitcode compilation and JIT-linking fix with no auth, crypto, or permission surfaces.

Level of scrutiny

Moderate. The patch_intrinsic pattern is used extensively in llvm_context.cpp for dozens of other intrinsics, so the approach is well-established. The one novel piece is the custom IR builder for cuda_match_all_sync_i32, which needs to handle the {i32, i1} struct return of nvvm_match_all_sync_i32p — the CreateExtractValue(result, {0}) call correctly pulls the i32 match value, matching the old inline-asm behavior of discarding the predicate.

Other factors

  • The author provided concrete crash reproduction and test output for both failure modes.
  • Existing test_active_mask, test_match_any, and test_match_all in test_simt.py cover the CUDA x86 path on CI T4 runners.
  • The pre-existing __linux___ typo (three underscores) flagged by the bug hunter is rendered moot by -mno-outline-atomics; the fix works regardless.
  • The pre-existing duplicate block_memfence call (flagged in a prior review) is harmless and out of scope for this PR.
  • The author addressed the cuda_match_any_sync_i64 concern by restoring the patch_intrinsic call with a comment explaining it IS called from C++ runtime code in node_pointer.h::is_representative().

Comment thread quadrants/runtime/llvm/runtime_module/runtime.cpp
@duburcqa
duburcqa force-pushed the duburcqa/fix_aarch64_invalid_asm branch from 9ddc9f2 to d33e7d2 Compare April 13, 2026 13:45

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — the inline-asm stubs are correctly replaced with LLVM NVPTX intrinsics, the dead AArch64 atomic helper block is removed, and -mno-outline-atomics cleanly solves the JIT linker issue.

Extended reasoning...

Overview

This PR touches three files: llvm_context.cpp (adds patch_intrinsic calls for cuda_active_mask, cuda_match_all_sync_i32, and cuda_match_any_sync_i64), CMakeLists.txt (adds -mno-outline-atomics to the clang command that compiles runtime bitcode), and runtime.cpp (replaces PTX inline-asm bodies with trivial stubs and removes the dead aarch64 atomic helper block).

Security risks

None. These are internal LLVM IR lowering helpers with no security surface.

Level of scrutiny

Moderate — this touches GPU codegen, but the patch_intrinsic pattern is already used for ~30 other intrinsics in the same file, making this a well-understood, low-risk extension of an existing mechanism. The custom IR wrapper for cuda_match_all_sync_i32 (ExtractValue from a struct return) is straightforward and correct. The -mno-outline-atomics flag is a no-op on x86 and documented with a clear rationale.

Other factors

Prior inline comments from this reviewer flagged (1) the duplicate block_memfence call (pre-existing, out of scope) and (2) the linux_ typo — the dead block has been fully removed in this PR, resolving that concern. The cuda_match_any_sync_i64 patch_intrinsic was re-added after the author confirmed it is live in node_pointer.h::is_representative(). CI tests (test_active_mask, test_match_any, test_match_all) cover the x86_64+CUDA path; the author provided direct test output demonstrating both crash modes on AArch64.

@duburcqa
duburcqa force-pushed the duburcqa/fix_aarch64_invalid_asm branch from d33e7d2 to b8c315e Compare April 13, 2026 14:05

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — the patch_intrinsic approach is correct and follows established patterns; the -mno-outline-atomics fix is well-justified.

Extended reasoning...

Overview

Three files changed: CMakeLists.txt adds -mno-outline-atomics to the bitcode compile flags; runtime.cpp removes dead PTX inline-asm blocks and replaces them with trivial stubs; llvm_context.cpp adds patch_intrinsic calls for cuda_active_mask, cuda_match_all_sync_i32 (with a custom IR builder to handle the {i32,i1} struct return), and cuda_match_any_sync_i64.

Security risks

None. This is GPU backend codegen, not auth or data-access code.

Level of scrutiny

Moderate — GPU codegen is non-trivial, but every change follows patterns already used dozens of times in the same file. The custom IR for match_all (CreateIntrinsic + CreateExtractValue) is straightforward. The author documented the rationale thoroughly and reproduced both crash modes. CI covers x86+CUDA via the T4 runner.

Other factors

The only flagged bug (CMakeLists ${arch} vs ${rtm_arch} in the install() line) is pre-existing and not touched by this PR. My two prior inline comments have been acknowledged. The author addressed all reviewer feedback (removed spurious cuda_match_any_sync_i64 patch_intrinsic, then correctly re-added it after finding it is called from node_pointer.h). No outstanding unresolved concerns.

Comment thread quadrants/runtime/llvm/runtime_module/CMakeLists.txt
@duburcqa

Copy link
Copy Markdown
Contributor Author

I was assisted by Claude Opus to write this PR. I have read and review every changes in this PR. I take full responsibility for the lines added and removed in this PR. I won't blame any issue on Claude Opus. I have reviewed this PR, and approve it.

@hughperkins hughperkins left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! 🙌

@hughperkins
hughperkins enabled auto-merge (squash) April 14, 2026 22:00
Comment on lines +10 to +11
# instead. It is a no-op on x86 and irrelevant for CUDA/GPU paths. The alternative (registering those symbols with
# the JIT linker) would be far more complex and fragile.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 The comment on CMakeLists.txt line 10 states that -mno-outline-atomics is 'irrelevant for CUDA/GPU paths', but this is factually incorrect — the flag is required for CUDA bitcode on AArch64 for the same reason it is required for CPU. The code itself is correct (the flag is applied to all architectures), but the inaccurate comment could mislead a future developer into removing the flag from CUDA-only builds, reintroducing an NVPTX backend crash; the phrase 'irrelevant for CUDA/GPU paths' should be dropped.

Extended reasoning...

The comment added by this PR reads: 'It is a no-op on x86 and irrelevant for CUDA/GPU paths.' While the first half is accurate, the second half is wrong in a way that creates a latent regression risk.

Why the flag is needed for CUDA bitcode on AArch64: runtime_cuda.bc is compiled from runtime.cpp on the HOST platform (AArch64) with -D ARCH_cuda. The file unconditionally includes atomic.h, which defines atomic_exchange_i32/atomic_exchange_u64 using atomic_exchange builtins (and similarly for atomic_and*, atomic_or, atomic_compare_exchange_ variants). On AArch64 without -mno-outline-atomics, clang lowers these GCC/Clang builtins to calls to __aarch64_swp4_acq_rel and __aarch64_swp8_acq_rel at the LLVM IR level, embedding them as undefined external references in the bitcode.

What patch_atomic_add does NOT cover: The patch_atomic_add lambda in module_from_file only patches atomic_add_i32, atomic_add_i64, atomic_add_f32, and atomic_add_f64. It leaves atomic_exchange_i32, atomic_exchange_u64, and all the atomic_and_/atomic_or_/atomic_xor_/atomic_min_/atomic_max_/atomic_compare_exchange_ variants completely unpatched. Without -mno-outline-atomics, these unpatched functions would retain __aarch64_swp* and __aarch64_ldadd* external calls in runtime_cuda.bc.

The crash path: module_from_file sets the CUDA module target triple to nvptx64-nvidia-cuda and links it with libdevice. Functions like mutex_lock_i32 and mutex_unlock_i32 (which call atomic_exchange_i32) are reachable from CUDA kernels via locked_task (used in allocate_from_reserved_memory, ListManager::touch_chunk, etc.), and ListManager::touch_chunk also calls atomic_exchange_u64. After eliminate_unused_functions retains anything reachable from runtime_* entry points, the __aarch64_swp* calls would survive into the NVPTX module. The NVPTX backend would then encounter unresolvable AArch64-specific external symbols and crash.

Step-by-step proof: (1) On AArch64, clang compiles atomic_exchange_i32 (from atomic.h) to an IR call to _aarch64_swp4_acq_rel when -moutline-atomics is active. (2) mutex_lock_i32 calls atomic_exchange_i32; locked_task calls mutex_lock_i32; allocate_from_reserved_memory calls locked_task; runtime_initialize calls allocate_from_reserved_memory. (3) runtime_initialize is a runtime* symbol, so it survives eliminate_unused_functions. (4) The __aarch64_swp4_acq_rel undefined reference thus survives into the NVPTX-targeted module. (5) The NVPTX backend rejects it and crashes.

Why the current code is actually correct: The -mno-outline-atomics flag IS applied uniformly to all architectures via the single add_custom_target command, so runtime_cuda.bc is compiled correctly with the flag present. The bug is purely in the comment, not the build logic.

The risk and fix: The inaccurate comment was introduced by this PR. A future developer reading 'irrelevant for CUDA/GPU paths' while refactoring the build might guard the flag with 'if(NOT CUDA_ARCH)', reintroducing the crash on any AArch64 + CUDA system. The fix is a one-line comment correction: remove the phrase 'and irrelevant for CUDA/GPU paths', leaving 'It is a no-op on x86.'

Addressing the refutation: The sole refutation agrees the comment is technically inaccurate but argues it is only a comment issue, not a code bug. That is correct — severity is nit, not normal. However, since this inaccurate comment was introduced by this PR and specifically describes the flag as safe to omit for GPU builds when it is not, it warrants a review note.

@hughperkins
hughperkins merged commit 5e8efa6 into main Apr 14, 2026
48 checks passed
@hughperkins
hughperkins deleted the duburcqa/fix_aarch64_invalid_asm branch April 14, 2026 23:03
npoulad1 added a commit to ROCm/quadrants that referenced this pull request Jun 8, 2026
* [Misc] Warn user to disable caching when print_ir/QD_DUMP_IR enabled (Genesis-Embodied-AI#425)

Co-authored-by: v01dxyz <v01dxyz@v01d.xyz>

* [Build] Pin torch version to CUDA 12.8 for CUDA tests (Genesis-Embodied-AI#428)

* [Misc] Fixing up taichi-dev urls (Genesis-Embodied-AI#429)

* [Perf] Rename cuda_graph to gpu_graph across the codebase (Genesis-Embodied-AI#430)

* Misc: fix typo integeral -> integral (Genesis-Embodied-AI#434)

Co-authored-by: v01dxyz <v01dxyz@v01d.xyz>

* [Perf] CUDA graph 4: call from multiple locations (Genesis-Embodied-AI#420)

* [Bug] Fix fastcache not restoring graph_do_while_arg (Genesis-Embodied-AI#435)

* [Perf] Cache last-call result in perf_dispatch for single-compatible case (Genesis-Embodied-AI#438)

* Fix gpu_graph fallback on old Nvidia GPU. (Genesis-Embodied-AI#443)

* Fix shared memory offset not reset between CUDA kernels. (Genesis-Embodied-AI#442)

* [Misc] Allow disabling GPU graph via QD_GPU_GRAPH=0 env var (Genesis-Embodied-AI#439)

* [Misc] Add named top-level loops (Genesis-Embodied-AI#440)

* [Misc] Rename gpu_graph to graph (Genesis-Embodied-AI#446)

* [Misc] Add cross-platform shuffle (Genesis-Embodied-AI#447)

* [Bug] Fix graph_do_while on Windows: search for cudadevrt.lib (Genesis-Embodied-AI#456)

* [Bug] Also search default CUDA toolkit install location on Windows (Genesis-Embodied-AI#461)

* [SPIRV] Feature Parity Atomics & Shared Array (Genesis-Embodied-AI#432)

* [Misc] Change clang format to 120 characters (Genesis-Embodied-AI#463)

* [Misc] CUDA graph 5 Add fatbin (Genesis-Embodied-AI#464)

* [Bug] Reuse VkInstance across init/reset cycles (Genesis-Embodied-AI#465)

* [Perf] Tiles 1: _load, _store, _eye_ (Genesis-Embodied-AI#466)

* [Misc] Remove dead InternalFuncStmt type_check override (Genesis-Embodied-AI#471)

* [Perf] Tiles 2: add cholesky and ger (Genesis-Embodied-AI#472)

* [Perf] Tiles 2b: add triangular solve (Genesis-Embodied-AI#474)

* [Misc] Refactor: use _get_col/_set_col in tiles load/store/init (Genesis-Embodied-AI#475)

* [Build] Fix flaky test_clock_accuracy (Genesis-Embodied-AI#436)

* Fix AARCH64 emitting invalid asm in CUDA kernels. (Genesis-Embodied-AI#473)

Co-authored-by: Hugh Perkins <hughperkins@gmail.com>

* [AMDGPU] Enable HIP memory pool and surface pool-exhaustion errors. (Genesis-Embodied-AI#485)

* [AMDGPU] Scope hsaco tmp dir per-user to avoid collisions. (Genesis-Embodied-AI#484)

* [Perf] Tiles 3: Add slice syntax, qd.outer() and initial doc (Genesis-Embodied-AI#477)

* [AMDGPU] Fix gradient computation. (Genesis-Embodied-AI#486)

* Enable all backends that are supported in unit tests. (Genesis-Embodied-AI#488)

* Fix SPIRV ID overflow for large kernels due to autodiff. (Genesis-Embodied-AI#489)

* [Misc] Fix purity checker to allow accessing constants from quadrants modules (Genesis-Embodied-AI#487)

* [Misc] Increase tolerance for clock monotonic test (Genesis-Embodied-AI#492)

* [CI] Serialize api doc workflow (Genesis-Embodied-AI#494)

* [CI] Increase tolerance for clock test (Genesis-Embodied-AI#506)

* [CI] Increase clock test tolerance to 20% (Genesis-Embodied-AI#509)

* [Perf] Add tensor_type parametrization to tile16 tests (Genesis-Embodied-AI#504)

* [Perf] Tiles 4b: Migrate tiles16 tests to enable fastcache (Genesis-Embodied-AI#505)

* [Perf] Tiles 4c: add Tiles16x16 proxy (Genesis-Embodied-AI#507)

* [Perf] Tiles 4d: Consolidate slice error tests using parametrize (Genesis-Embodied-AI#508)

* [Perf] Tiles 4: add SharedArray slice support (Genesis-Embodied-AI#482)

* [Perf] Tiles 5: add Cholesky benchmark demo (Genesis-Embodied-AI#483)

* [Doc] Add user guide page for subgroup shuffle (Genesis-Embodied-AI#512)

* [Perf] Implement cross-platform shuffle_down (Genesis-Embodied-AI#510)

* [Perf] Add portable subgroup reduce_add and reduce_all_add (Genesis-Embodied-AI#511)

* [Perf] Add first warmup config to perf dispatch (Genesis-Embodied-AI#422)

* [AutoDiff] Autodiff 1: Add baseline adstack regression test for unary_collections (Genesis-Embodied-AI#500)

* [AutoDiff] Autodiff 2: Implement derivative for tan (Genesis-Embodied-AI#501)

* [AutoDiff] Autodiff 3: Recompute tanh/exp on the operand in the reverse pass (Genesis-Embodied-AI#502)

* [AutoDiff] Autodiff 4: Mark rsqrt as non-linear for adstack promotion (Genesis-Embodied-AI#503)

* [AutoDiff] Autodiff 5: Fix adjoint-alloca placement for GlobalLoads outside the current range-for (Genesis-Embodied-AI#496)

* [AutoDiff] Autodiff 6: Adstack regression tests (Genesis-Embodied-AI#491)

* [AutoDiff] Autodiff 7: Fix header size in AdStackAllocaStmt to match u64 runtime layout (Genesis-Embodied-AI#534)

* [AutoDiff] Autodiff 8: Surface LLVM adstack push/pop overflow as a Python exception (Genesis-Embodied-AI#535)

* [AutoDiff] Autodiff 9: Guard against LLVM worker-thread stack overflow from large per-task adstack budget (Genesis-Embodied-AI#495)

* [AutoDiff] Autodiff 10: Implement adstack for SPIR-V (Genesis-Embodied-AI#490)

* [AutoDiff] Autodiff 11: Latent adstack-adjacent fixes (AMDGPU hipFree, flush() keeps ctx_buffers_, always-preallocate) (Genesis-Embodied-AI#536)

* [Doc] Add AGENTS.md with instructions for AI agents (Genesis-Embodied-AI#541)

* [Bug] Abort kernel execution on assertion failure instead of segfaulting (Genesis-Embodied-AI#419)

* [Type] ndarray typing 1: Add eval_str=True to inspect.signature() calls (Genesis-Embodied-AI#411)

* [CI] Suppress reportPrivateImportUsage in torch-using files (Genesis-Embodied-AI#552)

* [Misc] QD_DUMP_IR dumps to files with the task_id added to the filename (Genesis-Embodied-AI#441)

* [Type] ndarray typing 2: Fix NDArray single-arg subscript crash (Genesis-Embodied-AI#412)

* [Test] Flush xdist channel before worker exit so test failure reports are visible (Genesis-Embodied-AI#555)

* [CI] Reduce test retries on CI from 3 to 1. (Genesis-Embodied-AI#554)

* [AutoDiff] Autodiff 12: Heap-backed adstack on LLVM backends (CPU/CUDA/AMDGPU) (Genesis-Embodied-AI#537)

* [AutoDiff] Autodiff 13: Heap-backed adstack on SPIR-V backends (Metal, Vulkan) (Genesis-Embodied-AI#493)

* [AutoDiff] Autodiff 14: Resolve bounded-inner-loop adstacks without default_ad_stack_size fallback (Genesis-Embodied-AI#539)

* [SPIRV] Vulkan SPIR-V correctness: atomic-view aliasing, PSB stride, narrow storage caps, u1 cast, per-init layer recheck (Genesis-Embodied-AI#513)

* [Build] Autodiff 15: Replace 2022 MoltenVK pin with LunarG Vulkan SDK fetch and sanitise MoltenVK cap advertisement (Genesis-Embodied-AI#551)

* [Test] Suppress stock pytest-timeout to avoid conflict with pytest_hardtle (Genesis-Embodied-AI#557)

* [Vulkan] Use SDK validation layer for debugPrintf instead of apt package (Genesis-Embodied-AI#562)

* [Test] Fix flaky perf_dispatch tests by increasing work amounts (Genesis-Embodied-AI#559)

* [Test] Add --maxfail CLI option to run_tests.py (default 20) (Genesis-Embodied-AI#558)

* [CI] Vulkan debug printf fix to address flaky tests (Genesis-Embodied-AI#563)

* [Docs] Add a new page to help for first time contributors (Genesis-Embodied-AI#426)

Authored-by: v01dxyz <v01dxyz@v01d.xyz>

* [AutoDiff] Autodiff 16: Resolve reverse-mode adstack depths per-launch via runtime-evaluated SizeExpr (Genesis-Embodied-AI#543)

* Fix: raise error if device memory allocation fails (Genesis-Embodied-AI#451) (Genesis-Embodied-AI#453)

Co-authored-by: v01dxyz <v01dxyz@v01d.xyz>
Co-authored-by: Hugh Perkins <hughperkins@gmail.com>

* [CI] Add CI job to check line wrapping of comments and docs (Genesis-Embodied-AI#564)

* [Misc] Add coverage report to PRs, including kernels (Genesis-Embodied-AI#470)

* [CI] CI wrap check feeds only diffs to agent (Genesis-Embodied-AI#567)

* Skip 'flaky' test on MacOS CI. (Genesis-Embodied-AI#573)

* [Test] Fix missing `import sys` in test_fail_device_memory_allocation (Genesis-Embodied-AI#574)

* [CI] Fix Vulkan debugPrintf flake with session-scoped warmup (Genesis-Embodied-AI#571)

* [AutoDiff] determine_ad_stack_size: replace whole-CFG Bellman-Ford with SCC + DAG DP (Genesis-Embodied-AI#575)

* [Test] Fix macOS OOM skip reason to describe actual root cause (Genesis-Embodied-AI#576)

* [Lang] whole_kernel_cse: 2.5x compile time speedup on large kernels (Genesis-Embodied-AI#577)

* [CI] Add CI check for unnecessarily deleted comments (Genesis-Embodied-AI#570)

* [CI] Migrate coverage report to github Check page (Genesis-Embodied-AI#566)

* [Lang] Skip IR verifier between passes unless debug=true (Genesis-Embodied-AI#579)

* [Lang] Inline AdStack ops on release LLVM codegen: dramatically reduces compile time for adstack-enabled reverse-mode kernels (Genesis-Embodied-AI#584)

* [CUDA] Honor offline_cache=False end-to-end so QD_OFFLINE_CACHE=0 actually gives a cold compile (Genesis-Embodied-AI#580)

* [Type] Tensor 24 (Genesis-Embodied-AI#561)

Co-authored-by: hugh <hugh@slurm-login-0.slurm-login.tenant-slurm.svc.cluster.local>

* [Lang] auto_diff host-walk reductions: dramatically faster front-end compile time on adstack-enabled reverse-mode kernels (Genesis-Embodied-AI#587)

* [AutoDiff] Speed up reverse-mode kernel launches on GPU backends (Genesis-Embodied-AI#578)

* [Vulkan] Move adstack-sizer scratch out of Function-scope memory to fix SPIR-V pipeline build failures (Genesis-Embodied-AI#588)

* [AutoDiff] Improve diagnosis of unsupported reverse-mode AD patterns (Genesis-Embodied-AI#590)

* [Bug] Fix: promote Ndarray to AnyArray in build_Name for flattened struct fields (Genesis-Embodied-AI#592)

* [SPIR-V] Shrink reverse-grad kernel MSL by ~50% (Genesis-Embodied-AI#591)

* [CI] Add CI check that PR changes have test coverage (Genesis-Embodied-AI#596)

* [Perf] Enable zero-copy in to_torch() and to_numpy() (Genesis-Embodied-AI#450)

* Add BufferView: safe sub-range ndarray access for kernels (Genesis-Embodied-AI#585)

Co-authored-by: alanray-tech <alanray-tech@users.noreply.github.com>
Co-authored-by: Hugh Perkins <hughperkins@gmail.com>

* [Doc] Add user-facing fastcache documentation (Genesis-Embodied-AI#597)

Co-authored-by: hugh <hugh@slurm-login-0.slurm-login.tenant-slurm.svc.cluster.local>

* [Misc] Upgrade to enable v1 dlpack so to_numpy(copy=False) writable (Genesis-Embodied-AI#598)

Co-authored-by: root <root@rtx-209-201.slurm-compute.tenant-slurm.svc.cluster.local>

* [AutoDiff] Cut reverse-mode adstack memory usage 10x on all backends (Genesis-Embodied-AI#599)

* [Misc] Add CI check for feature file factorization (Genesis-Embodied-AI#606)

* [Perf] Skip _recursive_set_args for all-Field frozen dataclass structs (Genesis-Embodied-AI#607)

Co-authored-by: Cursor <cursoragent@cursor.com>

* [AutoDiff] SNode-arm bound-expr capture rejects fold-attack gate indices (Genesis-Embodied-AI#610)

* [Misc] Suppress field fastcache warning for qd.Tensor (Genesis-Embodied-AI#615)

Co-authored-by: Cursor <cursoragent@cursor.com>

* [AutoDiff] Adstack heap: clip reducer count by per-task loop trip count (compile-time and SizeExpr-evaluated) (Genesis-Embodied-AI#611)

* [Misc] Forward copy= through qd.Tensor, add copy=None option (Genesis-Embodied-AI#616)

Co-authored-by: Cursor <cursoragent@cursor.com>

* [Doc] Update README (Genesis-Embodied-AI#617)

Co-authored-by: Cursor <cursoragent@cursor.com>

* [CI] Fix coverage report showing def lines as uncovered (Genesis-Embodied-AI#623)

Co-authored-by: Cursor <cursoragent@cursor.com>

* [Perf] Generic launcher: persistent context, JIT-pointer reuse, Metal compute encoder, LLVM-GPU async memory ops (Part 1/2) (Genesis-Embodied-AI#619)

* [CI] Encode Python-first testing policy in coverage-check prompt (Genesis-Embodied-AI#622)

Co-authored-by: Cursor <cursoragent@cursor.com>

* [CI] Add PR Line change report (Genesis-Embodied-AI#624)

Co-authored-by: Cursor <cursoragent@cursor.com>

* [CI] Disable quadrants pytest plugin during quadrants internal coverage runs (Genesis-Embodied-AI#629)

Co-authored-by: Cursor <cursoragent@cursor.com>

* [AutoDiff] Adstack load+store eliminations: EliminateRecomputableAdStackPushes pass + leaf extensions (Genesis-Embodied-AI#621)

* [CI] Simplify coverage PR comment to a single linked line (Genesis-Embodied-AI#630)

* [CUDA] Add AGX Thor, SM_110 (Genesis-Embodied-AI#631)

Co-authored-by: Johnny Nunez and Hugh Perkins

* [CI] Lines changed report: collapse PR comment to a single linked totals line (Genesis-Embodied-AI#632)

* [FEATURE] Support external Metal command queue via qd.init (Genesis-Embodied-AI#618)

Co-authored-by: Cursor <cursoragent@cursor.com>

* [Perf] Cache adstack-sizer metadata per task across SPIR-V + LLVM-GPU; per-snode / DeviceAllocation invalidation (Part 2/2) (Genesis-Embodied-AI#620)

* [AutoDiff] Disable EliminateRecomputableAdStackPushes pending mutated-SNode chain-leaf fix (Genesis-Embodied-AI#633)

* [AutoDiff] Adstack chain-clone safety: mutated-SNode leaf reject + load_top consumer-aware guard (Genesis-Embodied-AI#634)

* [Docs] Add user-guide page for qd.simt.block.* primitives (Genesis-Embodied-AI#638)

* [Docs] Expand qd.simt.subgroup user-guide page to cover every op (Genesis-Embodied-AI#639)

* [Perf] Streams 1-4 (Genesis-Embodied-AI#410)

* [Docs] Add user-guide page for matrix decompositions and solvers (Genesis-Embodied-AI#643)

* [Bug] Revert "[Perf] Streams 1-4 (Genesis-Embodied-AI#410)" (Genesis-Embodied-AI#650)

* [Docs] Add user-guide page for atomics and bit operations (Genesis-Embodied-AI#640)

* [Docs] Add user-guide page for qd.simt.grid.* primitives (Genesis-Embodied-AI#641)

* [AutoDiff] Adstack max-reducer: parallel multi-axis MaxOverRange dispatch (Genesis-Embodied-AI#635)

* [AMDGPU] Fix amdgpu parallel rand init (Genesis-Embodied-AI#658)

* [Perf] Adstack: skip max-reducer recognizer on CPU + lift host-eval cap (Genesis-Embodied-AI#655)

* [Perf] Re-land Streams 1-4 with bug fixes (Genesis-Embodied-AI#653)

* [AMDGPU] Apply device_memory_GB=0.3 cap to AMDGPU tests (Genesis-Embodied-AI#659)

* [Perf] Per-launch host sync: drop wait_idle on SPIR-V, pin stream and drop stream_synchronize on CUDA/AMDGPU (Genesis-Embodied-AI#654)

* [AMDGPU] Unload hipModule_t in JITModuleAMDGPU destructor (Genesis-Embodied-AI#660)

* [AMDGPU] Trim default mempool on qd.reset() (Genesis-Embodied-AI#669)

* [AMDGPU] Hoist rand-state buffer to process lifetime (Genesis-Embodied-AI#668)

* [Streams] Use events for streams serialization on AMDGPU and CUDA (Genesis-Embodied-AI#667)

* [Perf] Adstack max-reducer: launch cache + zero-copy result map; content-stable registry_id (Genesis-Embodied-AI#671)

* [SPIR-V] dispatch_max_reducers: register each task with the real kernel name (Genesis-Embodied-AI#675)

* [AutoDiff] Debug-mode field/grad/dual: dtype, layout, and access-time invariants (Genesis-Embodied-AI#677)

* [Docs] Add user-guide page for qd.algorithms.* device-wide algorithms (Genesis-Embodied-AI#642)

Co-authored-by: alanray-tech <alan.ray@genesis-ai.company>

* [Docs] Doc for existing atomics: switch support table to per-backend columns (Genesis-Embodied-AI#657)

Co-authored-by: alanray-tech <alan.ray@genesis-ai.company>

* [GPU] Cross gpu atomics (Genesis-Embodied-AI#666)

Co-authored-by: alanray-tech <alan.ray@genesis-ai.company>

* [GPU] Make block operations portable cross-gpu (Genesis-Embodied-AI#664)

* [Perf] CPU LLVM adstack-cache: skip per-launch bump-writes + ndarray_shapes capture on forward-only handles (Genesis-Embodied-AI#685)

* [GPU] Cross-GPU for grid ops (Genesis-Embodied-AI#670)

* [Math] Make bitop operations portable cross-gpu (Genesis-Embodied-AI#662)

* [AMDGPU] Always use wave64, on both RDNA and CDNA (Genesis-Embodied-AI#687)

* [AMDGPU] Use syncscope("agent") for atomix xor to avoid CAS livelock (Genesis-Embodied-AI#672)

* [GPU] New bit ops for QIPC (Genesis-Embodied-AI#679)

* [GPU] Subgroup ops cross-gpu (Genesis-Embodied-AI#665)

* [Graph] Rename CUDA Graph to Graph in docs (Genesis-Embodied-AI#691)

* [SPIR-V] Fix FIFO-queue ordering when sharing command queue. (Genesis-Embodied-AI#694)

* [Atomics] New QIPC ops for atomics (Genesis-Embodied-AI#690)

* Pass dataclass sub-structs into qd.func (Genesis-Embodied-AI#698)

* [AMDGPU] HIP graph runtime support for @qd.kernel(graph=True) (Genesis-Embodied-AI#692)

* [CI] Add per-file timing report to Mac Metal test job (Genesis-Embodied-AI#695)

Co-authored-by: Cursor <cursoragent@cursor.com>

* [CI] Enable kernel disk cache during tests (Genesis-Embodied-AI#696)

* [Math] New QIPC ops for single-threaded linalg (Genesis-Embodied-AI#683)

* [BREAKING][GPU] New QIPC ops for subgroups (Genesis-Embodied-AI#676)

* [GPU] New QIPC ops for block (Genesis-Embodied-AI#684)

* [GPU] New device-level ops for QIPC (Genesis-Embodied-AI#693)

* [algorithms] PrefixSumExecutor: drop unused GRID_SZ local (Genesis-Embodied-AI#701)

* [block] sync(): fix unsupported-arch error message (Genesis-Embodied-AI#700)

* [volatile_load] add qd.volatile_load primitive (closes Genesis-Embodied-AI#648) (Genesis-Embodied-AI#702)

* [AutoDiff] Reject recycled identity_key in AdStackCache::register_adstack_sizing_info (Genesis-Embodied-AI#708)

* [Vulkan] Declare GroupNonUniform SPIR-V caps and enable shaderSubgroupExtendedTypes (Genesis-Embodied-AI#707)

* Fix duplicate HIP graph driver-function declarations after v1.0.0 merge

The amd-integration fork had cherry-picked the HIP graph driver functions
(graph_create / graph_destroy / graph_add_kernel_node / graph_instantiate /
graph_exec_destroy / graph_launch), and upstream v1.0.0 added the same set.
The per-file 3-way merge appended both copies into
amdgpu_driver_functions.inc.h, producing redeclaration errors that broke the
AMDGPU RHI/runtime compile. Drop the upstream duplicate block; the signatures
are identical to the fork's existing declarations.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix AMDGPU launcher coherence and num_instructions visibility after v1.0.0 merge

- kernel_launcher.cpp: the 3-way merge spliced upstream v1.0.0's launch_llvm_kernel
  rewrite (ephemeral arg/context buffers, explicit-stream path, AmdgpuDefaultStream
  PinGuard) onto the AMD fork's kernarg-by-value + persistent-scratch design,
  leaving references to undefined `ephemeral_context_ptr`. Restore the fork's
  coherent launch_llvm_kernel verbatim; it calls the (already merged) enhanced
  launch_offloaded_tasks, which keeps the max-reducer dispatch and stream-parallel
  groups adapted onto the AMD launch path.
- llvm_context.h: both the fork and upstream added `num_instructions`; the merge
  kept upstream's private placement, but the AMDGPU codegen force-inline heuristic
  calls it statically from outside the class. Move it back to the public section.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Restore async result D2H and hoist kernarg vectors in AMDGPU launcher

The v1.0.0 merge resolution regressed two amd-integration baseline
optimizations in launch_llvm_kernel / launch_offloaded_tasks:

  - The per-launch result-buffer copy was a blocking memcpy_device_to_host,
    forcing a host stall on every value-returning launch and serializing the
    GPU pipeline. Restore the async D2H (the caller synchronizes lazily when it
    needs the value); external-array transfers still stream_synchronize once
    before reading back.

  - launch_task constructed the kernarg std::vectors from initializer lists
    ({kernarg_payload} / {kernarg_size}) on every dispatch (heap alloc + free
    per launch). Hoist arg_ptrs/arg_sizes out of the per-task launch and reuse.

Co-authored-by: Cursor <cursoragent@cursor.com>

* amdgpu: default to LDS permlane64 emulation; drop host-x86 barrier asm on retarget

Two AMDGPU JIT-compile crashes surfaced after the v1.0.0 merge pulled in the QIPC subgroup
ops (Genesis-Embodied-AI#676), which made the rigid constraint solver's wave-cooperative reductions route through
`amdgpu_cross_half_shuffle_i32`. Both manifested as a SIGSEGV inside
`llvm::SIInstrInfo::getInstSizeInBytes` during `JITSessionAMDGPU::compile_module_to_hsaco`
(i.e. at first kernel launch), and reproduce on gfx942 / MI300X. Baseline 0.4.6 never emitted
these constructs, which is why it was unaffected.

1. Native `llvm.amdgcn.permlane64` lowering crashes the bundled LLVM 22.1.0 AMDGPU backend.
   Default `amdgpu_permlane64` to the existing LDS-roundtrip software emulation on every target
   (it produces identical results). Add `QD_AMDGPU_USE_NATIVE_PERMLANE64=1` to opt back into the
   native instruction once the backend bug is fixed; the old `QD_AMDGPU_FORCE_PERMLANE64_FALLBACK`
   is now the default and still honored. This is the actual crash fix.

2. The runtime module is compiled by the host x86_64 clang and only retargeted to amdgcn here, so
   `amdgpu_cross_half_shuffle_i32`'s `__asm__ volatile("" : "+v"(byte))` optimization barrier carries
   x86 flag clobbers (`~{dirflag},~{fpsr},~{flags}`) that are meaningless on AMDGPU. The IR verifies
   but the empty-body INLINEASM is invalid on the amdgcn target. Neutralize empty-body barrier asm
   during retarget (forward the tied value, then erase) so no stale host asm reaches codegen. On the
   wave64 targets we ship `ds_bpermute` already addresses the full wave, so the hint is a no-op.

Co-authored-by: Cursor <cursoragent@cursor.com>

* style: apply clang-format (v19.1.7) to AMDGPU fn_attrs and launcher sources

CI pre-commit's clang-format hook reformatted these files (long
declarations/lambda signatures collapsed onto single lines per the repo's
clang-format config). Apply the same formatting so the hook passes.

No functional changes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(amdgpu): use CreateNeg for branchless i32 sgn instead of CreateSub(0, input)

clang-tidy (modernize-use-nullptr, -warnings-as-errors) flagged
`builder->CreateSub(0, input)` in the i32 sgn path: the literal `0` binds to
the `llvm::Value*` LHS parameter as a null pointer, not an integer zero.
Replace with `builder->CreateNeg(input)`, which emits `0 - input` with a proper
zero constant -- identical intended semantics, and clang-tidy clean.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Robert Dazi <14996868+v01dXYZ@users.noreply.github.com>
Co-authored-by: v01dxyz <v01dxyz@v01d.xyz>
Co-authored-by: Hugh Perkins <hughperkins@gmail.com>
Co-authored-by: Alexis DUBURCQ <alexis.duburcq@gmail.com>
Co-authored-by: hugh <hugh@slurm-login-0.slurm-login.tenant-slurm.svc.cluster.local>
Co-authored-by: alanray-tech <alan.ray@genesis-ai.company>
Co-authored-by: alanray-tech <alanray-tech@users.noreply.github.com>
Co-authored-by: root <root@rtx-209-201.slurm-compute.tenant-slurm.svc.cluster.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Johnny <johnnynuca14@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants