Skip to content

target_features: sse (or at least avx2) is incompatible with soft-float ABI - #160302

Open
RalfJung wants to merge 2 commits into
rust-lang:mainfrom
RalfJung:soft-float-no-avx2
Open

target_features: sse (or at least avx2) is incompatible with soft-float ABI#160302
RalfJung wants to merge 2 commits into
rust-lang:mainfrom
RalfJung:soft-float-no-avx2

Conversation

@RalfJung

@RalfJung RalfJung commented Jul 31, 2026

Copy link
Copy Markdown
Member

Fixes #117938

Enabling both the avx2 and soft-float target features is not supported by LLVM and can crash the backend. Let's preempt that with rust-level checks. (I still think there's also an LLVM bug here, it shouldn't just SIGILL on unexpected target feature configurations, but that's a different discussion.)

What is not clear to me is whether this just affects just avx2 or also avx or even sse (we don't support mmx/3dnow separately). @dianqk do you know more about this? To be safe, let's reject "sse" and therefore by implication also all other x86 vector target features.

This PR turns #[target_feature(enable = "sse")] on a softfloat target into an FCW similar to what we do on aarch64 (see #135160). The FCW only affects people building for soft-float targets which is a fairly small percentage of our overall users (and which means we cannot meaningfully crater this). I hence went for "report in deps" immediately so that the people building the actual binaries see these warnings that their upstreams probably will never see.

@rustbot rustbot added S-waiting-on-review Status: Awaiting review from the assignee but also interested parties. T-compiler Relevant to the compiler team, which will review and decide on the PR/issue. labels Jul 31, 2026
@rustbot

rustbot commented Jul 31, 2026

Copy link
Copy Markdown
Collaborator

r? @khyperia

rustbot has assigned @khyperia.
They will have a look at your PR within the next two weeks and either review your PR or reassign to another reviewer.

Use r? to explicitly pick a reviewer

Why was this reviewer chosen?

The reviewer was selected based on:

  • Owners of files modified in this PR: compiler
  • compiler expanded to 75 candidates
  • Random selection from 16 candidates

@RalfJung RalfJung changed the title target_featurs: avx2 is incompatible with soft-float ABI target_features: avx2 is incompatible with soft-float ABI Jul 31, 2026
@RalfJung

Copy link
Copy Markdown
Member Author

r? @workingjubilee or @dianqk

@rustbot

rustbot commented Jul 31, 2026

Copy link
Copy Markdown
Collaborator

workingjubilee is currently at their maximum review capacity.
They may take a while to respond.

@RalfJung

Copy link
Copy Markdown
Member Author

Hm unfortunately it seems like LLVM does not crash on all functions with #[target_feature(enable = "avx2")]. This one for example works fine:

#[unsafe(no_mangle)]
#[target_feature(enable = "avx2")]
pub fn foobar(x: __m256i, y: __m256i) -> __m256i {
    _mm256_or_si256(x, y)
}

So we may have to add an FCW for this after all.

@tarcieri do you know which function is causing the trouble in dalek-cryptography/curve25519-dalek#601? All we know is that it's somewhere in poly1305...

@RalfJung

Copy link
Copy Markdown
Member Author

Okay I have a reproducer:

#![no_std]

use core::arch::x86_64::*;

#[unsafe(no_mangle)]
#[target_feature(enable = "avx2")]
pub fn foobar(ptr: *const __m256i) -> __m256i { unsafe {
    let key = _mm256_loadu_si256(ptr);
    _mm256_and_si256(
        _mm256_permutevar8x32_epi32(key, _mm256_set_epi32(3, 7, 2, 6, 1, 5, 0, 4)),
        _mm256_set_epi32(0, -1, 0, -1, 0, -1, 0, -1),
    )
}}

@RalfJung
RalfJung force-pushed the soft-float-no-avx2 branch 2 times, most recently from 58062d1 to 21a23e9 Compare July 31, 2026 20:47
@RalfJung RalfJung added the I-lang-nominated Nominated for discussion during a lang team meeting. label Jul 31, 2026
@RalfJung RalfJung changed the title target_features: avx2 is incompatible with soft-float ABI target_features: sse (or at least avx2) is incompatible with soft-float ABI Jul 31, 2026
@dianqk

dianqk commented Aug 1, 2026

Copy link
Copy Markdown
Member

I don't know these features on x86, but the PR seems reasonable to me.

@RalfJung

RalfJung commented Aug 1, 2026

Copy link
Copy Markdown
Member Author

With Nikita on vacation, who might know which features LLVM supports on x86 in combination with +soft-float?

But I guess we can also just warn about all vector features (as this PR does now) and if we get issues saying sse actually works fine we can always adjust. 🤷

@RalfJung

RalfJung commented Aug 1, 2026

Copy link
Copy Markdown
Member Author

FWIW the s390x target also has a "soft-float" target feature and there we already mark "vector" as incompatible.

ARM also has "soft-float" and there we don't mark anything. ARM also has much more explicit ABI control so maybe setting FloatABIType to "soft" and enabling neon actually works fine there? No idea.

@rust-bors

This comment has been minimized.

@RalfJung
RalfJung force-pushed the soft-float-no-avx2 branch from 21a23e9 to 5594310 Compare August 5, 2026 06:30
@rustbot

rustbot commented Aug 5, 2026

Copy link
Copy Markdown
Collaborator

This PR was rebased onto a different main commit. Here's a range-diff highlighting what actually changed.

Rebasing is a normal part of keeping PRs up to date, so no action is needed—this note is just to help reviewers.

@traviscross traviscross added I-lang-radar Items that are on lang's radar and will need eventual work or consideration. P-lang-drag-1 Lang team prioritization drag level 1. https://rust-lang.zulipchat.com/#narrow/channel/410516-t-lang labels Aug 5, 2026
Comment thread compiler/rustc_lint_defs/src/builtin.rs
@traviscross traviscross added the T-lang Relevant to the language team label Aug 5, 2026
@traviscross

Copy link
Copy Markdown
Contributor

Makes sense to me. Thanks @RalfJung.

@rfcbot fcp merge lang

@rust-rfcbot

rust-rfcbot commented Aug 5, 2026

Copy link
Copy Markdown
Collaborator

@traviscross has proposed to merge this. The next step is review by the rest of the tagged team members:

No concerns currently listed.

Once a majority of reviewers approve (and at most 2 approvals are outstanding), this will enter its final comment period. If you spot a major issue that hasn't been raised at any point in this process, please speak up!

cc @rust-lang/lang-advisors: FCP proposed for lang, please feel free to register concerns.
See this document for info about what commands tagged team members can give me.

@rust-rfcbot rust-rfcbot added proposed-final-comment-period Proposed to merge/close by relevant subteam, see T-<team> label. Will enter FCP once signed off. disposition-merge This issue / PR is in PFCP or FCP with a disposition to merge it. labels Aug 5, 2026
@scottmcm

scottmcm commented Aug 5, 2026

Copy link
Copy Markdown
Member

Definitely happy to just say "no, you can't do that" for strange combinations.

@RalfJung

RalfJung commented Aug 5, 2026

Copy link
Copy Markdown
Member Author

Given that this might just crash the backend anyway,

It might crash, but it does not always crash. I haven't found any documentation on the LLVM side about which target feature combinations they support and which they don't support.

Definitely happy to just say "no, you can't do that" for strange combinations.

The combination isn't that strange IMO -- you can use a soft-float ABI but still locally have the instruction & registers available. But LLVM just doesn't (always) support that.

@Ranger3143

Copy link
Copy Markdown

A real-world data point in support of this FCW, from a target you can't crater — and a failure mode that isn't SIGILL.

I hit this on x86_64-unknown-uefi, whose target spec carries +soft-float:

$ rustc --print target-spec-json --target x86_64-unknown-uefi   # (nightly, -Zunstable-options)
"features": "-mmx,-sse,+soft-float"

The crate is a no_std inference engine with hand-written AVX2 kernels — #[target_feature(enable = "avx2")] on the hot functions, _mm256_* intrinsics throughout, built --release. An objdump census of the binary we had actually shipped:

staged binary (stock x86_64-unknown-uefi), 453,120 bytes
  census: xmm=0  ymm=0  vfmadd=0

Zero vector instructions. Every f32 op, including inside the target_feature(enable = "avx2") bodies, was lowered to soft-float libcalls.

It did not crash. No SIGILL, no link error, no warning. The binary booted and produced correct output — the soft-float path preserves single-rounding semantics — it was just catastrophically slow. Measured against the same workload on the same machine (Dell Inspiron 15, i5-5200U, bare metal), rebuilding against a custom target spec with the float features removed took one hot function from 1.65–1.67 × 10⁹ ticks/token to 6.48–7.59 × 10⁶, a 217–258× recovery, and the resulting binary was ~50 KB smaller:

hardfloat rebuild, 402,432 bytes
  census: xmm=6855  ymm=433  vfmadd=210

We spent weeks attributing this to our own code. The only thing that found it was counting instructions in the shipped artifact; we now gate staging on a census that fails when ymm == 0.

So: strong +1 on warning rather than staying silent. The SIGILL cases at least announce themselves. This one ships.

Scope-wise I can only speak to avx2 on x86_64-unknown-uefi; I haven't tested plain sse or avx in this configuration.

Full write-up and logs, if useful: https://github.com/Aefinity-AI/alice-aegis/blob/main/docs/posts/2026-08-05_uefi-soft-float-deletes-your-avx2.md

@RalfJung

RalfJung commented Aug 5, 2026

Copy link
Copy Markdown
Member Author

Thanks for the feedback!

(Note that under our just-adopted AI policy, we consider it improper to just copy-paste AI reports into PR/issue comments. Please use your own words and put the detailed AI report into a clearly indicated, separate section that one can easily skip.)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

disposition-merge This issue / PR is in PFCP or FCP with a disposition to merge it. I-lang-nominated Nominated for discussion during a lang team meeting. I-lang-radar Items that are on lang's radar and will need eventual work or consideration. P-lang-drag-1 Lang team prioritization drag level 1. https://rust-lang.zulipchat.com/#narrow/channel/410516-t-lang proposed-final-comment-period Proposed to merge/close by relevant subteam, see T-<team> label. Will enter FCP once signed off. S-waiting-on-review Status: Awaiting review from the assignee but also interested parties. T-compiler Relevant to the compiler team, which will review and decide on the PR/issue. T-lang Relevant to the language team

Projects

None yet

Development

Successfully merging this pull request may close these issues.

LLVM produces SIGILL when enabling avx2 target feature on x86_64-unknown-none

9 participants