Skip to content

[MLAS] fix: add runtime check for AVX512-FP16 support instead of relying on compiler version - #28236

Open
PriscillaJCorn wants to merge 1 commit into
microsoft:mainfrom
PriscillaJCorn:main
Open

[MLAS] fix: add runtime check for AVX512-FP16 support instead of relying on compiler version#28236
PriscillaJCorn wants to merge 1 commit into
microsoft:mainfrom
PriscillaJCorn:main

Conversation

@PriscillaJCorn

Copy link
Copy Markdown

Description

It is inappropriate to merely guess whether the AVX512-FP16 instruction is supported based on the compiler's version information, as this can lead to numerous exceptions. Therefore, before starting, I have added a simple and practical check, which involves testing the compiler's support for the instruction through a small case study.

Motivation and Context

Fix #22519 and #24025
The current logic determines AVX512-FP16 support based on compiler version heuristics.
This is not reliable, since compiler version alone does not guarantee that the target
toolchain or hardware fully supports the instruction set.

As a result, this may lead to incorrect feature enablement, causing build failures or
runtime exceptions on certain environments.

This change improves robustness by replacing the heuristic with a compile-time check,
ensuring that AVX512-FP16 is only enabled when it is actually supported.

@PriscillaJCorn

Copy link
Copy Markdown
Author

@PriscillaJCorn please read the following Contributor License Agreement(CLA). If you agree with the CLA, please reply with the following information.

@microsoft-github-policy-service agree [company="{your company}"]

Options:

  • (default - no company specified) I have sole ownership of intellectual property rights to my Submissions and I am not making Submissions in the course of work for my employer.
@microsoft-github-policy-service agree
  • (when company given) I am making Submissions in the course of work for my employer (or my employer has intellectual property rights in my Submissions by contract or applicable law). I have permission from my employer to make Submissions and enter into this Agreement on behalf of my employer. By signing below, the defined term “You” includes me and my employer.
@microsoft-github-policy-service agree company="Microsoft"

Contributor License Agreement

Contribution License Agreement

This Contribution License Agreement (“Agreement”) is agreed to by the party signing below (“You”), and conveys certain license rights to Microsoft Corporation and its affiliates (“Microsoft”) for Your contributions to Microsoft open source projects. This Agreement is effective as of the latest signature date below.

  1. Definitions.
    “Code” means the computer software code, whether in human-readable or machine-executable form,
    that is delivered by You to Microsoft under this Agreement.
    “Project” means any of the projects owned or managed by Microsoft and offered under a license
    approved by the Open Source Initiative (www.opensource.org).
    “Submit” is the act of uploading, submitting, transmitting, or distributing code or other content to any
    Project, including but not limited to communication on electronic mailing lists, source code control
    systems, and issue tracking systems that are managed by, or on behalf of, the Project for the purpose of
    discussing and improving that Project, but excluding communication that is conspicuously marked or
    otherwise designated in writing by You as “Not a Submission.”
    “Submission” means the Code and any other copyrightable material Submitted by You, including any
    associated comments and documentation.
  2. Your Submission. You must agree to the terms of this Agreement before making a Submission to any
    Project. This Agreement covers any and all Submissions that You, now or in the future (except as
    described in Section 4 below), Submit to any Project.
  3. Originality of Work. You represent that each of Your Submissions is entirely Your original work.
    Should You wish to Submit materials that are not Your original work, You may Submit them separately
    to the Project if You (a) retain all copyright and license information that was in the materials as You
    received them, (b) in the description accompanying Your Submission, include the phrase “Submission
    containing materials of a third party:” followed by the names of the third party and any licenses or other
    restrictions of which You are aware, and (c) follow any other instructions in the Project’s written
    guidelines concerning Submissions.
  4. Your Employer. References to “employer” in this Agreement include Your employer or anyone else
    for whom You are acting in making Your Submission, e.g. as a contractor, vendor, or agent. If Your
    Submission is made in the course of Your work for an employer or Your employer has intellectual
    property rights in Your Submission by contract or applicable law, You must secure permission from Your
    employer to make the Submission before signing this Agreement. In that case, the term “You” in this
    Agreement will refer to You and the employer collectively. If You change employers in the future and
    desire to Submit additional Submissions for the new employer, then You agree to sign a new Agreement
    and secure permission from the new employer before Submitting those Submissions.
  5. Licenses.
  • Copyright License. You grant Microsoft, and those who receive the Submission directly or
    indirectly from Microsoft, a perpetual, worldwide, non-exclusive, royalty-free, irrevocable license in the
    Submission to reproduce, prepare derivative works of, publicly display, publicly perform, and distribute
    the Submission and such derivative works, and to sublicense any or all of the foregoing rights to third
    parties.
  • Patent License. You grant Microsoft, and those who receive the Submission directly or
    indirectly from Microsoft, a perpetual, worldwide, non-exclusive, royalty-free, irrevocable license under
    Your patent claims that are necessarily infringed by the Submission or the combination of the
    Submission with the Project to which it was Submitted to make, have made, use, offer to sell, sell and
    import or otherwise dispose of the Submission alone or with the Project.
  • Other Rights Reserved. Each party reserves all rights not expressly granted in this Agreement.
    No additional licenses or rights whatsoever (including, without limitation, any implied licenses) are
    granted by implication, exhaustion, estoppel or otherwise.
  1. Representations and Warranties. You represent that You are legally entitled to grant the above
    licenses. You represent that each of Your Submissions is entirely Your original work (except as You may
    have disclosed under Section 3). You represent that You have secured permission from Your employer to
    make the Submission in cases where Your Submission is made in the course of Your work for Your
    employer or Your employer has intellectual property rights in Your Submission by contract or applicable
    law. If You are signing this Agreement on behalf of Your employer, You represent and warrant that You
    have the necessary authority to bind the listed employer to the obligations contained in this Agreement.
    You are not expected to provide support for Your Submission, unless You choose to do so. UNLESS
    REQUIRED BY APPLICABLE LAW OR AGREED TO IN WRITING, AND EXCEPT FOR THE WARRANTIES
    EXPRESSLY STATED IN SECTIONS 3, 4, AND 6, THE SUBMISSION PROVIDED UNDER THIS AGREEMENT IS
    PROVIDED WITHOUT WARRANTY OF ANY KIND, INCLUDING, BUT NOT LIMITED TO, ANY WARRANTY OF
    NONINFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE.
  2. Notice to Microsoft. You agree to notify Microsoft in writing of any facts or circumstances of which
    You later become aware that would make Your representations in this Agreement inaccurate in any
    respect.
  3. Information about Submissions. You agree that contributions to Projects and information about
    contributions may be maintained indefinitely and disclosed publicly, including Your name and other
    information that You submit with Your Submission.
  4. Governing Law/Jurisdiction. This Agreement is governed by the laws of the State of Washington, and
    the parties consent to exclusive jurisdiction and venue in the federal courts sitting in King County,
    Washington, unless no federal subject matter jurisdiction exists, in which case the parties consent to
    exclusive jurisdiction and venue in the Superior Court of King County, Washington. The parties waive all
    defenses of lack of personal jurisdiction and forum non-conveniens.
  5. Entire Agreement/Assignment. This Agreement is the entire agreement between the parties, and
    supersedes any and all prior agreements, understandings or communications, written or oral, between
    the parties relating to the subject matter hereof. This Agreement may be assigned by Microsoft.

@microsoft-github-policy-service agree

@hariharans29

Copy link
Copy Markdown
Member

/azp run Linux QNN CI Pipeline,Win_TRT_Minimal_CUDA_Test_CI,Windows ARM64 QNN CI Pipeline,Windows GPU Doc Gen CI Pipeline

@azure-pipelines

Copy link
Copy Markdown
No pipelines are associated with this pull request.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR improves MLAS AVX512-FP16 feature enablement by replacing compiler-version heuristics with an actual compile-time capability check, so AVX512-FP16 codepaths are only enabled when the toolchain can assemble the required instructions (addressing build failures reported in #22519 and #24025).

Changes:

  • Replace _MSC_VER/__GNUC__ version heuristics in MLAS runtime dispatch with a build-defined capability macro (MLAS_SUPPORTS_AVX512FP16).
  • Add a CMake compile test for AVX512-FP16 instruction support and conditionally include the AVX512-FP16 conversion assembly.
  • Minor whitespace cleanup in ARM FP16 kernel selection code.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 4 comments.

File Description
onnxruntime/core/mlas/lib/platform.cpp Switches runtime AVX512-FP16 dispatch gating to a build-time macro instead of compiler-version checks.
cmake/onnxruntime_mlas.cmake Adds a compile-time probe for AVX512-FP16 instruction support and conditionally enables related assembly + a preprocessor define.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines 529 to +537
#ifndef __APPLE__
#if (defined(_MSC_VER) && (_MSC_VER >= 1933)) || (defined(__GNUC__) && (__GNUC__ >= 13))
#if defined(MLAS_SUPPORTS_AVX512FP16)
//
// Check if the processor supports AVX NE CONVERT.
//
if ((Cpuid7_1[3] & (0b1 << 5)) != 0) {
this->CastF16ToF32Kernel = &MlasCastF16ToF32KernelAvx;
}
#endif // (defined(_MSC_VER) && (_MSC_VER >= 1933)) || (defined(__GNUC__) && (__GNUC__ >= 13))
#endif // MLAS_SUPPORTS_AVX512FP16

Copilot AI Apr 27, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

MLAS_SUPPORTS_AVX512FP16 now gates enabling CastF16ToF32Kernel, but that macro is only defined in the non-MSVC CMake path. On Windows/MSVC, setup_mlas_source_for_windows() can still add amd64/cvtfp16Avx.asm (MSVC_VERSION>=1933), yet this block will never run, so AVX512-FP16 support may be silently disabled at runtime even when the kernel is built. Consider defining MLAS_SUPPORTS_AVX512FP16 in the MSVC path when cvtfp16Avx.asm is enabled (or reintroduce an MSVC-specific compile-time guard).

Copilot uses AI. Check for mistakes.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@copilot apply changes based on this feedback Yes, in func setup_mlas_source_for_windows() , we should define MLAS_SUPPORTS_AVX512FP16 in the MSVC path when cvtfp16Avx.asm is enabled.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this comment addressed ?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this comment addressed ?

Yes !

Comment thread cmake/onnxruntime_mlas.cmake Outdated
Comment thread cmake/onnxruntime_mlas.cmake Outdated
Comment thread cmake/onnxruntime_mlas.cmake Outdated
@hariharans29

Copy link
Copy Markdown
Member

Thanks for the contribution - can you please check if copilot's review comments are relevant ?

@hariharans29

Copy link
Copy Markdown
Member

/azp run Linux QNN CI Pipeline,Win_TRT_Minimal_CUDA_Test_CI,Windows ARM64 QNN CI Pipeline,Windows GPU Doc Gen CI Pipeline

@azure-pipelines

Copy link
Copy Markdown
No pipelines are associated with this pull request.

@PriscillaJCorn

Copy link
Copy Markdown
Author

/azp run Linux QNN CI Pipeline,Win_TRT_Minimal_CUDA_Test_CI,Windows ARM64 QNN CI Pipeline,Windows GPU Doc Gen CI Pipeline
@hariharans29 Please check and re-runthis CI. It seems that some packages failed to download during the build process.

@hariharans29
hariharans29 requested a review from Copilot April 29, 2026 17:18

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated 3 comments.


💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread cmake/onnxruntime_mlas.cmake Outdated
Comment on lines 799 to 800
list(APPEND mlas_private_compile_definitions MLAS_SUPPORTS_AVX512FP16)
endif()

Copilot AI Apr 29, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The probe is hardwired to GCC/Clang-style flags (-mavx512fp16) and GNU inline asm syntax, which will fail under MSVC (and some clang-cl configurations). This can silently disable the FP16 path on Windows even when the toolchain supports it, changing behavior vs the prior _MSC_VER branch. Consider branching on CMAKE_CXX_COMPILER_ID and using a C++-based probe with compiler-appropriate flags (e.g., MSVC /arch: equivalent or a clang-cl compatible approach), so the macro is correctly set across supported toolchains.

Copilot uses AI. Check for mistakes.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@copilot apply changes based on this feedback

Comment thread cmake/onnxruntime_mlas.cmake Outdated
)
if(CMAKE_CXX_COMPILER_VERSION GREATER_EQUAL 13.1 AND NOT(APPLE))

include(CheckCSourceCompiles)

Copilot AI Apr 29, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This feature is consumed by C++ sources (platform.cpp), but the probe uses CheckCSourceCompiles. Using CheckCXXSourceCompiles (and the corresponding include) makes the probe match the actual compilation environment (C++ compiler, flags, and frontend), and avoids coupling this logic to the presence/configuration of the C language in the build.

Copilot uses AI. Check for mistakes.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@copilot apply changes based on this feedback

Comment on lines +530 to +537
#if defined(MLAS_SUPPORTS_AVX512FP16)
//
// Check if the processor supports AVX NE CONVERT.
//
if ((Cpuid7_1[3] & (0b1 << 5)) != 0) {
this->CastF16ToF32Kernel = &MlasCastF16ToF32KernelAvx;
}
#endif // (defined(_MSC_VER) && (_MSC_VER >= 1933)) || (defined(__GNUC__) && (__GNUC__ >= 13))
#endif // MLAS_SUPPORTS_AVX512FP16

Copilot AI Apr 29, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The compile-time guard name MLAS_SUPPORTS_AVX512FP16 doesn’t match the runtime check/comment (“AVX NE CONVERT”) and the instruction used in the CMake probe (vcvtneeph2ps). Please align terminology so it’s clear what’s actually being required (e.g., rename the macro/probe variable to reflect AVX-NE-CONVERT, or update the comment/probe to match AVX512-FP16 if that’s the intended feature).

Copilot uses AI. Check for mistakes.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@copilot apply changes based on this feedback

@hariharans29 hariharans29 changed the title fix: add runtime check for AVX512-FP16 support instead of relying on compiler version [MLAS] fix: add runtime check for AVX512-FP16 support instead of relying on compiler version May 19, 2026
@JingliangGao

Copy link
Copy Markdown

@PriscillaJCorn Thinkl you, this patch help me solve the compilation problem. Will this patch be merged to the main branch? I believe this patch will help many people.

@PriscillaJCorn

Copy link
Copy Markdown
Author

@PriscillaJCorn Thinkl you, this patch help me solve the compilation problem. Will this patch be merged to the main branch? I believe this patch will help many people.

I have been very busy with work recently and haven't paid attention to this issue for a long time. Umn, I also think this PR is very interesting. Wait a moment, I will modify the PR and resubmit it.

@PriscillaJCorn

Copy link
Copy Markdown
Author

@hariharans29 With @JingliangGao 's generous contribution, we have resolved the issue and hope you can review this PR again and consider accepting it.

@hariharans29

Copy link
Copy Markdown
Member

Review: PR #28236 — [MLAS] fix: add runtime check for AVX512-FP16 support instead of relying on compiler version

Request changes — the intent is right (replace compiler-version heuristics with a real capability probe), but the probe as written is very likely to compile-fail on every toolchain and silently disable the kernel it's meant to fix. A few naming/scoping issues also make it hard to reason about correctness. All are addressable with small edits.

Blocking

  1. The probe intrinsic name looks wrong and will make check_c_source_compiles always return false. Both branches probe with:

    __m256 v = _mm256_cvtneephi2ps(_mm256_setzero_si256());

    I could not find _mm256_cvtneephi2ps in the Intel Intrinsics Guide, in GCC's avxneconvertintrin.h, or in Clang/MSVC headers. The AVX-NE-CONVERT intrinsics that map to vcvtneeph2ps are _mm256_cvtneeph_ps (even) and _mm256_cvtneoph_ps (odd), and they take __m128h (a pointer-loaded half vector) as input, not __m256i. So the current probe fails at name-lookup and at argument-type checking on every compiler I can think of — which means COMPILER_SUPPORTS_AVX512FP16 will be FALSE universally, cvtfp16Avx.{asm,S} is never added to the sources, MLAS_SUPPORTS_AVX512FP16 is never defined, and CastF16ToF32Kernel = &MlasCastF16ToF32KernelAvx is never enabled. Net effect: this PR silently regresses AVX-NE-CONVERT on the exact toolchains (MSVC ≥ 1933, GCC ≥ 13) where the old heuristic used to enable it. Please compile this probe locally against MSVC and GCC 13+ and paste the CMake configure log showing COMPILER_SUPPORTS_AVX512FP16 = TRUE before we can merge. Suggested body that actually resolves through the intrinsic headers on both compilers:

    #include <immintrin.h>
    int main(void) {
        __m128i x = _mm_setzero_si128();       // 8 fp16 lanes as bits
        __m256 y  = _mm256_cvtneeph_ps((const __m128h*)&x);
        (void)y;
        return 0;
    }

    (or use _mm256_cvtneoph_ps — either resolves iff the AVX-NE-CONVERT intrinsics are visible.)

  2. Wrong -m flag on the GCC/Clang branch. The non-MSVC branch adds -mavx512fp16 to CMAKE_REQUIRED_FLAGS. AVX-512-FP16 (Sapphire Rapids Xeon) and AVX-NE-CONVERT (Alder Lake E-core, Sapphire Rapids, and everything newer) are separate ISA extensions. The kernel that gets compiled (x86_64/cvtfp16Avx.S) uses vcvtneeph2ps, which is AVX-NE-CONVERT. The correct GCC/Clang flag is -mavxneconvert (GCC 12+, Clang 14+). Compiling the probe with -mavx512fp16 on GCC 13 will still fail to expose _mm256_cvtneeph_ps because the intrinsic is header-gated on __AVXNECONVERT__, not __AVX512FP16__. Same story for the runtime bit you're checking (Cpuid7_1[3] & (1 << 5) = AVX-NE-CONVERT bit, not FP16). Please switch to -mavxneconvert.

  3. The MSVC branch probably doesn't need /arch:AVX512 — and adding it can break the parent build. MSVC exposes AVX-NE-CONVERT intrinsics under its normal AVX2 headers on recent versions; you don't need /arch:AVX512 to compile a probe that uses _mm256_cvtneeph_ps. More importantly, /arch:AVX512 in CMAKE_REQUIRED_FLAGS will unconditionally influence the compile line for check_c_source_compiles, and if the developer's default is /arch:AVX2 (typical for this project), a probe-succeeded result can arise from an environment that the rest of the MLAS build never actually compiles with. Drop /arch:AVX512 for the MSVC probe, or replace it with a scoped try_compile that mirrors the real MLAS compile flags.

  4. The macro name still lies about what it gates. MLAS_SUPPORTS_AVX512FP16 gates:

    • a runtime CPUID.(EAX=7,ECX=1):EDX.bit5 check that is AVX-NE-CONVERT, per Intel SDM Vol. 2A;
    • inclusion of cvtfp16Avx.{asm,S} which uses AVX-NE-CONVERT mnemonics (vcvtneeph2ps);
    • a CMake probe that (once fixed per items 1–2) will detect AVX-NE-CONVERT.

    None of the three things is AVX-512-FP16. Please rename MLAS_SUPPORTS_AVX512FP16MLAS_SUPPORTS_AVX_NECONVERT (and rename COMPILER_SUPPORTS_AVX512FP16 correspondingly; the // Check if the processor supports AVX NE CONVERT. comment is already correct — keep it). Otherwise the next person to touch this code will read the macro name and add real AVX-512-FP16 gating under it, breaking non-Sapphire-Rapids Alder Lake targets that only have AVX-NE-CONVERT.

Non-blocking

  1. Use check_cxx_source_compiles. platform.cpp is C++, the actual onnxruntime_mlas target is compiled as C++, and any per-target C++ flags (e.g. -std=c++17, custom -march= from parent scope) won't be seen by the C-compiler probe. Include CheckCXXSourceCompiles instead of CheckCSourceCompiles.
  2. Duplication. The 12-line "save CMAKE_REQUIRED_FLAGS → set → probe → restore" block is copy-pasted between setup_mlas_source_for_windows() and the non-MSVC else-branch, with only the added flag differing. Factor into a helper macro (_mlas_check_avxneconvert_support(<flag>)) so we don't drift again.
  3. Whitespace-only churn. The non-MSVC hunk includes ~6 lines of pure re-indentation of the existing Using -mavx2 -mfma -mavxvnni message/setter block. Please strip those from the diff — they make the review noisy and will conflict with any unrelated edits in that block.
  4. PARENT_SCOPE bubble-up is fine on the Windows side (setup_mlas_source_for_windows is a function, so PARENT_SCOPE on mlas_private_compile_definitions at the end is needed and correctly added). Confirm the non-MSVC branch that also appends to mlas_private_compile_definitions is at file scope (not inside a function). From the diff it appears to be inside a top-level else(), so no PARENT_SCOPE needed — but double-check the enclosing scope has no function(...) wrapper introduced in the meantime.
  5. CI evidence. Please rerun full CI on the current HEAD and confirm at least one Linux-x64 leg produces a build log line showing COMPILER_SUPPORTS_AVX512FP16 (or its renamed version) set to TRUE. Silently building without the kernel would look like green CI but is exactly the failure mode this PR aims to avoid.
  6. Motivation-vs-fix framing in the PR description. The description says the change replaces a heuristic with "a compile-time check, ensuring that AVX512-FP16 is only enabled when it is actually supported." The runtime CPUID check for Cpuid7_1[3] & (1<<5) is preserved from before, so the CPU-side gate hasn't changed — the fix is only on the compile-time side, ensuring the assembler/compiler can emit the required opcodes. Worth clarifying in the PR body so reviewers don't look for a runtime-detection change that isn't there.
  7. File name. cvtfp16Avx.{asm,S} is also misleadingly named — it's an AVX-NE-CONVERT kernel, not AVX-{2,512}. Rename is out of scope for this PR but worth a follow-up.

Summary

Two focused edits — (a) fix the probe body to use a real AVX-NE-CONVERT intrinsic, (b) switch the GCC/Clang flag to -mavxneconvert — should get this over the line. Rename the macro at the same time to prevent future confusion, and please post the configure log on both MSVC and GCC 13+ showing the probe succeeded. Happy to re-review as soon as those are in.

@hariharans29 hariharans29 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for the PR. Please take a look at the pasted review comment. If you could please address them, we can merge it

@blazingphoenix7

Copy link
Copy Markdown
Contributor

Flagging a functional problem with the compile probe itself, separate from the naming point above. _mm256_cvtneephi2ps is not a valid intrinsic, so check_c_source_compiles never succeeds and MLAS_SUPPORTS_AVX512FP16 is left undefined on every toolchain. The build still passes, but CastF16ToF32Kernel is then never assigned in platform.cpp, which silently drops the AVX-NE-CONVERT fp16 cast that the old _MSC_VER >= 1933 / __GNUC__ >= 13 guard used to enable, so the change reverses its own intent.

Reproduced locally:

  • gcc 15 and clang 21, with -mavx512fp16, -mavxneconvert, or both: call to undeclared function '_mm256_cvtneephi2ps'
  • MSVC /arch:AVX512: same, the name does not resolve (it falls through to an implicit int and then fails with C2440)

The instruction the kernel uses is vcvtneeph2ps, which is AVX-NE-CONVERT rather than AVX512-FP16. Its intrinsic is _mm256_cvtneeph_ps (both gcc 15 and clang 21 recognize that name under -mavxneconvert, and note it takes a memory operand, not an __m256i). So the probe needs the correct intrinsic with a matching -mavxneconvert flag, or the inline-asm form suggested earlier in the thread. As written it can only ever evaluate to false.

@blazingphoenix7

Copy link
Copy Markdown
Contributor

Following up with the root cause and a probe I verified, in case it saves some time.

Why the compiler version check fails in the first place: the failures in #24025 and #22519 happen at assembly time, and the assembler moves independently of the compiler. The #24025 reporter had g++ 13.1.0 with binutils 2.38, so the current CMAKE_CXX_COMPILER_VERSION GREATER_EQUAL 13.1 gate passes while as still rejects vcvtneeph2ps, and #22519 is the same shape with amdclang's integrated assembler. So what is worth probing is whether the assembler accepts the two mnemonics that cvtfp16Avx.S and cvtfp16Avx.asm actually contain, rather than whether the C compiler knows an intrinsic.

For the GCC and Clang path this evaluates correctly, following the inline asm direction suggested earlier in the thread:

check_c_source_compiles("
int main(void) {
    __asm__ __volatile__(\"vcvtneeph2ps (%rdi), %ymm0\");
    __asm__ __volatile__(\"vcvtneoph2ps (%rdi), %ymm1\");
    return 0;
}
" MLAS_ASSEMBLER_SUPPORTS_AVX_NE_CONVERT)

Run through cmake with gcc 15 and clang 21 it sets the variable to 1, while the _mm256_cvtneephi2ps version comes back empty on both. No ISA flag is needed for the probe or for the source: cvtfp16Avx.S assembles as-is with binutils 2.46 under the existing -mavx2 -mfma -mf16c -mavxvnni flags, so nothing needs to move out of mlas_platform_srcs_avx2 for this to work.

For MSVC, cvtfp16Avx.asm is assembled by ml64 rather than cl, so a check_c_source_compiles probe cannot gate it, and the GNU inline asm above will not compile there either. enable_language(ASM_MASM) already runs in adjust_global_compile_flags.cmake, so CMAKE_ASM_MASM_COMPILER is available:

set(NE_PROBE_SRC "${CMAKE_BINARY_DIR}/ne_probe.asm")
file(WRITE "${NE_PROBE_SRC}"
".code\n"
"probe PROC\n"
"    vcvtneeph2ps ymm0, ymmword ptr [rcx]\n"
"    vcvtneoph2ps ymm1, ymmword ptr [rcx]\n"
"    ret\n"
"probe ENDP\n"
"END\n")
execute_process(
  COMMAND "${CMAKE_ASM_MASM_COMPILER}" /nologo /c /Fo "${CMAKE_BINARY_DIR}/ne_probe.obj" "${NE_PROBE_SRC}"
  RESULT_VARIABLE NE_RV OUTPUT_QUIET ERROR_QUIET)

That returns 0 on ml64 14.44 here, and it also gives the MSVC path somewhere to define the macro, which was the first review comment on platform.cpp.

On naming, the header of cvtfp16Avx.asm describes the module as AVX_NE_CONVERT and these are AVX-NE-CONVERT instructions rather than AVX512-FP16, so something like MLAS_ASSEMBLER_SUPPORTS_AVX_NE_CONVERT would match what is actually being tested and line up with the earlier terminology comment.

Happy to send this as a PR against your branch if that is easier.

JingliangGao added a commit to PriscillaJCorn/onnxruntime that referenced this pull request Jul 30, 2026
…icrosoft#28236 review)

- Fix probe intrinsic name: _mm256_cvtneephi2ps -> _mm256_cvtneeph_ps
  with correct __m128h input type per Intel Intrinsics Guide
- Switch GCC/Clang flag from -mavx512fp16 to -mavxneconvert
  (AVX-NE-CONVERT and AVX-512-FP16 are separate ISA extensions)
- Remove unnecessary /arch:AVX512 from MSVC probe
  (MSVC exposes AVX-NE-CONVERT intrinsics under normal AVX2 headers)
- Rename macros to reflect actual capability:
  COMPILER_SUPPORTS_AVX512FP16 -> COMPILER_SUPPORTS_AVX_NECONVERT
  MLAS_SUPPORTS_AVX512FP16 -> MLAS_SUPPORTS_AVX_NECONVERT
- Use check_cxx_source_compiles instead of check_c_source_compiles
  (MLAS is a C++ target, C++ flags need to be respected)

Fixes review points from @hariharans29 in PR microsoft#28236
…icrosoft#28236 review)

- Fix probe intrinsic name: _mm256_cvtneephi2ps -> _mm256_cvtneeph_ps
  with correct __m128h input type per Intel Intrinsics Guide
- Switch GCC/Clang flag from -mavx512fp16 to -mavxneconvert
  (AVX-NE-CONVERT and AVX-512-FP16 are separate ISA extensions)
- Remove unnecessary /arch:AVX512 from MSVC probe
  (MSVC exposes these intrinsics under normal AVX2 headers on recent versions)
- Rename macros to reflect actual capability:
  COMPILER_SUPPORTS_AVX512FP16 -> COMPILER_SUPPORTS_AVX_NECONVERT
  MLAS_SUPPORTS_AVX512FP16 -> MLAS_SUPPORTS_AVX_NECONVERT
- Use check_cxx_source_compiles instead of check_c_source_compiles
  (MLAS is a C++ target, C++ flags need to be respected)

Fixes review points from @hariharans29 in PR microsoft#28236
@PriscillaJCorn

Copy link
Copy Markdown
Author

Hi, @hariharans29. Thanks so much for the super detailed review.

Um... I have listened to your opinions and made some modifications. In short, it can be reflected in the following aspects:

  • Intrinsic name_mm256_cvtneephi2ps was indeed wrong, replaced with _mm256_cvtneeph_ps((const __m128h*)&x) using your suggested probe body
  • GCC/Clang flag-mavxneconvert (you are right, they are separate ISA extensions)
  • MSVC probe — removed /arch:AVX512 since not needed
  • Macro renameAVX_NECONVERT everywhere to match what we actually detect
  • Probe language — switched to C++ probe (check_cxx_source_compiles)

BTW, sorry about the first pass, I will test locally before next time.

@blazingphoenix7 detection now correctly checks AVX-NE-CONVERT support instead of silently failing on all toolchains.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Build] compilation error: invalid instruction mnemonic 'vcvtneeph2ps'

5 participants