Skip to content

Set up CI with Azure Pipelines - #1

Closed
azure-pipelines[bot] wants to merge 1 commit into
masterfrom
azure-pipelines
Closed

Set up CI with Azure Pipelines#1
azure-pipelines[bot] wants to merge 1 commit into
masterfrom
azure-pipelines

Conversation

@azure-pipelines

Copy link
Copy Markdown

No description provided.

@msftclas

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission, we really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.

❌ azure-pipelines[bot] sign now
You have signed the CLA already but the status is still pending? Let us recheck it.

@raymondxyang
raymondxyang deleted the azure-pipelines branch November 19, 2018 20:22
tmccrmck added a commit to tmccrmck/onnxruntime that referenced this pull request Aug 28, 2019
groszewn pushed a commit to groszewn/onnxruntime that referenced this pull request Nov 13, 2019
…essor_double

Update TreeEnsembleRegressor type support
scxiao referenced this pull request in ROCm/onnxruntime Feb 4, 2020
titaiwangms added a commit to titaiwangms/onnxruntime that referenced this pull request May 6, 2026
Tianlei BLOCKER microsoft#1: New mode-1+softcap differentiating test (C++ + Python).
  With softcap > 0 active, qk_matmul_output_mode=1 (post-microsoft#7913 numbering =
  kPostSoftCap) snapshots softcap*tanh(scale*QK/softcap) with NO mask added.
  Without softcap, mode 1 aliases mode 0, so the swap is observationally
  indistinguishable — this test is what proves the 1<->2 swap actually
  changed semantics correctly.

Tianlei BLOCKER microsoft#2: New softcap+nonpad_kv_seqlen leakage test (C++ + Python).
  Exercises the latent fix where the nonpad sentinel is now applied AFTER
  softcap (per onnx#7867 ordering). Pre-fix: tanh squashed the sentinel,
  leaking poison V at padded positions through softmax.

Bot inline minors:
  - microsoft#3 (test_gqa.py): clarify fp16 docstring — CPU does support fp16; fp32 is
    the natural EP-native dtype for the canary.
  - microsoft#4 (attention_op_test.cc): regen comment now cites shared opset 23/24
    ordering and notes RunTest4D builds at opset 23.
  - microsoft#5 (attention_parameters.h): typo defintion -> definition.
  - microsoft#6 (attention.cc): replace 'guaranteed -inf' with precise wording citing
    mask_filter_value<T>() = numeric_limits::lowest() / MLFloat16::MinValue
    sentinel and the MLAS softmax finite-input requirement (attention.h).

R-2 microsoft#1 (attention_parameters.h): Spec-leading documentation block on the
  QKMatMulOutputMode enum noting that ORT now uses the post-onnx#7913
  numbering, while the bundled cmake/external/onnx (v1.21.0) still reflects
  the old numbering. ORT leads the spec change pending the next bundled-ONNX
  bump.

Plumbing: common.py attention_prompt_func gains an optional output_qk
  kwarg (default 0 / disabled). When > 0, returns a 4-tuple including the
  qk_matmul snapshot tensor; otherwise unchanged 3-tuple. No existing callers
  are affected.

Test results:
  - AttentionTest.* — 60/60 PASS (was 58, +2 new).
  - TestONNXAttentionCPUSoftcapMaskOrdering — 4/4 PASS (was 2, +2 new).
  - lintrunner clean across all 5 touched files.

Refs: lead-39245992/upstream-pr-status-recheck.md, pr1v2-review-{code,critical,readability,qa}.md

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
chilo-ms pushed a commit that referenced this pull request Jun 1, 2026
### Description

Add an internal session config entry, `"session.compile_only"`, set by
`CompileModel()` before
  session initialization. The NvTensorRTRTX EP reads it in
`NvExecutionProviderInfo::FromProviderOptions()` and, when set, skips
`deserializeCudaEngine()` /
  `createExecutionContext()` in `CreateNodeComputeInfoFromGraph()`.

The EP context node is still saved — that path uses the serialized
engine buffer directly and does
not depend on the deserialized engine. A stub compute function is
registered to satisfy the
framework; it returns `NOT_IMPLEMENTED` if called, which cannot happen
in practice because
  compile-only sessions are destroyed without inference.


### Motivation and Context

`OrtCompileAPI::CompileModel()` creates an `InferenceSession` solely to
drive `EP::Compile()` and
write out the EPContext model, then destroys it without running
inference. During that session, the
NvTensorRTRTX EP was performing a full `deserializeCudaEngine()` and
`createExecutionContext()` —
uploading engine weights to the GPU and JIT-ing the engine, only to free
everything when the session
  was destroyed.

When the user then loads the EPContext model in a real session, the same
JIT and upload happen again.
   Net effect on the typical "compile, then load and run" flow:

  ```
  ONNX model
      → CompileModel()         [JIT + GPU upload #1 — discarded]
      → EP context model saved to disk
      → Session from EP context model
                               [JIT + GPU upload #2 — necessary]
      → Inference
  ```

  JIT and GPU upload run twice.
tianleiwu added a commit that referenced this pull request Jul 27, 2026
)

### Description

`import onnxruntime` segfaults during `dlopen` of
`onnxruntime_pybind11_state.so` on Linux. This removes the global
initializer in `onnxruntime_pybind_state.cc` that eagerly calls
`Env::Default()`, and resolves the platform `Env` on first use at its
two call sites instead.

### Motivation and Context

The module has a namespace-scope dynamic initializer:

```cpp
static Env& platform_env = Env::Default();
```

Since POSIX telemetry landed (#27379), `Env::Default()` constructs
`PosixEnv`, whose `PosixTelemetry` member initializes the 1DS SDK in its
constructor. That path reads `defaultRuntimeConfig`, a namespace-scope
`static ILogConfiguration` defined in the 1DS SDK's
`RuntimeConfig_Default.hpp`, which lives in a **different translation
unit of the same shared library**.

Dynamic initialization order across translation units is unspecified,
and the pybind TU's initializer runs first. `Variant::merge_map`
therefore iterates a still zero-initialized `std::map`: `_M_node_count
== 0`, but `_M_header._M_left` is `nullptr` rather than self-pointing,
so `begin() != end()` and the loop dereferences null.

Textbook static initialization order fiasco. Backtrace (Release build
relinked without `--strip-all` to recover symbols):

```
#0  std::map<..., Variant>::lower_bound        stl_map.h:1307
#1  std::map<..., Variant>::operator[]         stl_map.h:509
#2  Variant::merge_map                         VariantType.hpp:508
#3  RuntimeConfig_Default::RuntimeConfig_Default  RuntimeConfig_Default.hpp:97
#4  LogManagerImpl::LogManagerImpl             LogManagerImpl.cpp:183
#5  LogManagerFactory::Create                  LogManagerFactory.cpp:36
#6  LogManagerFactory::lease
#7  LogManagerFactory::Get                     LogManagerFactory.hpp:71
#8  LogManagerProvider::Get                    LogManagerProvider.cpp:16
#9  onnxruntime::PosixTelemetry::Initialize()
#10 onnxruntime::PosixTelemetry::PosixTelemetry()
#11 onnxruntime::(anonymous namespace)::PosixEnv::PosixEnv()
#12 onnxruntime::Env::Default()
#13 _GLOBAL__sub_I_onnxruntime_pybind_state.cc
#14 call_init                                  elf/dl-init.c:74
...
#21 _dl_open                                   elf/dl-open.c:905
```

The crash is independent of `LD_LIBRARY_PATH` and `CUDA_VISIBLE_DEVICES`
— it happens before any ORT runtime code runs. `libonnxruntime.so` is
unaffected because it contains no global initializer that reaches
`Env::Default()`, which is why C/C++ and onnxruntime-genai consumers do
not see it.

Both remaining uses of `platform_env` are inside pybind lambdas that run
long after load, so calling `Env::Default()` there is safe. This also
removes the now-stale `TODO: we may delay-init this variable` and a
pre-existing `#pragma warning(push)` that should have been `pop`.

Note for follow-up: `Env::Default()` is now unsafe to call from any
dynamic initializer. This was the only such call site in the tree, but
hardening `PosixTelemetry` to defer SDK initialization out of its
constructor would remove the hazard entirely.

### Tests

Verified on a Linux CUDA 13 Release build (`onnxruntime_USE_TELEMETRY`
enabled):

- `dlopen` of `onnxruntime_pybind11_state.so` succeeds (previously
SIGSEGV).
- `import onnxruntime` reports the version and
`['CUDAExecutionProvider', 'CPUExecutionProvider']`.
- `onnxruntime.enable_telemetry_events()` / `disable_telemetry_events()`
— the two call sites changed here — work.
- CPU and CUDA inference sessions produce correct results.
- Reproduced and verified with `LD_LIBRARY_PATH` unset and
`CUDA_VISIBLE_DEVICES` empty.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants