Add debug messages for failure - #3
Closed
suryasidd wants to merge 0 commit into
Closed
Conversation
MaajidKhan
pushed a commit
that referenced
this pull request
Mar 29, 2021
* Test re-using page layout from current ONNX Runtime website for docs * Add content for documentation on website * Fixed most broken links * Copy just-the-docs theme sources into repo * Remove local theme files as this did not work with GitHub * Remove nojekyll file * Move image assets into single location * Add Contents to markdown files and ensure only one h1 * Update after review * Fix img links * Add trailing slash to main nav links * Fix broken links on main docs page * Re-fix broken links on main docs page * Fix broken links #3 * Fix broken links #4 * Fix broken links #5 * Fix broken links #6 * Fix paths to global assets * Add updates since fork * Update custom op docs * Fix link
sfatimar
pushed a commit
that referenced
this pull request
Dec 8, 2021
…de if sparse tensors are disabled. (microsoft#9898) * Add 2 builds to validate the cmake defines for excluding optional components work in both full and minimal builds. * Create empty config for no-ops build * Create empty config for no-ops build - attempt #2 * Create empty config for no-ops build - attempt #3 * Update python binding code to work when sparse tensors are disabled.
MaajidKhan
pushed a commit
that referenced
this pull request
Feb 15, 2022
…icrosoft#10260) * update java API for STVM EP. Issue is from PR#10019 * use_stvm -> use_tvm * rename stvm worktree * STVMAllocator -> TVMAllocator * StvmExecutionProviderInfo -> TvmExecutionProviderInfo * stvm -> tvm for cpu_targets. resolve onnxruntime::tvm and origin tvm namespaces conflict * STVMRunner -> TVMRunner * StvmExecutionProvider -> TvmExecutionProvider * tvm::env_vars * StvmProviderFactory -> TvmProviderFactory * rename factory funcs * StvmCPUDataTransfer -> TvmCPUDataTransfer * small clean * STVMFuncState -> TVMFuncState * USE_TVM -> NUPHAR_USE_TVM * USE_STVM -> USE_TVM * python API: providers.stvm -> providers.tvm. clean TVM_EP.md * clean build scripts #1 * clean build scripts, java frontend and others #2 * once more clean #3 * fix build of nuphar tvm test * final transfer stvm namespace to onnxruntime::tvm * rename stvm->tvm * NUPHAR_USE_TVM -> USE_NUPHAR_TVM * small fixes for correct CI tests * clean after rebase. Last renaming stvm to tvm, separate TVM and Nuphar in cmake and build files * update CUDA support for TVM EP * roll back CudaNN home check * ERROR for not positive input shape dimension instead of WARNING * update documentation for CUDA * small corrections after review * update GPU description * update GPU description * misprints were fixed * cleaned up error msgs Co-authored-by: Valery Chernov <valery.chernov@deelvin.com> Co-authored-by: KJlaccHoeUM9l <wotpricol@mail.ru> Co-authored-by: Thierry Moreau <tmoreau@octoml.ai>
saurabhkale17
pushed a commit
that referenced
this pull request
Jul 6, 2023
Fix memory leak issue which comes from TRT EP's allocator object not
being released upon destruction.
Following is the log from valgrind:
```
==1911860== 100,272 (56 direct, 100,216 indirect) bytes in 1 blocks are definitely lost in loss record 1,751 of 1,832
==1911860== at 0x483CFA3: operator new(unsigned long) (vg_replace_malloc.c:472)
==1911860== by 0x315DC2: std::_MakeUniq<onnxruntime::OrtAllocatorImplWrappingIAllocator>::__single_object std::make_unique<onnxruntime::OrtAllocatorImplWrappingIAllocator, std::shared_ptr<onnxruntime::IAllocator> >(std::shared_ptr<onnxruntime::IAllocator>&&) (unique_ptr.h:857)
==1911860== by 0x30EE7B: OrtApis::KernelContext_GetAllocator(OrtKernelContext const*, OrtMemoryInfo const*, OrtAllocator**) (custom_ops.cc:121)
==1911860== by 0x660D115: onnxruntime::TensorrtExecutionProvider::Compile(std::vector<onnxruntime::IExecutionProvider::FusedNodeAndGraph, std::allocator<onnxruntime::IExecutionProvider::FusedNodeAndGraph> > const&, std::vector<onnxruntime::NodeComputeInfo, std::allocator<onnxruntime::NodeComputeInfo> >&)::{lambda(void*, OrtApi const*, OrtKernelContext*)#3}::operator()(void*, OrtApi const*, OrtKernelContext*) const (tensorrt_execution_provider.cc:2223)
```
This issue happens after this [EP allocator
refactor](microsoft#15833)
sfatimar
pushed a commit
that referenced
this pull request
Aug 28, 2023
### Description
Release OrtEnv before main function returns. Before this change, OrtEnv
is deleted when C/C++ runtime destructs all global variables in ONNX
Runtime's core framework.
The callstack is like this:
```
* frame #0: 0x00007fffee39f5a6 libonnxruntime.so.1.16.0`onnxruntime::Environment::~Environment(this=0x00007fffee39fbf2) at environment.h:20:7
frame #1: 0x00007fffee39f614 libonnxruntime.so.1.16.0`std::default_delete<onnxruntime::Environment>::operator()(this=0x00007ffff4c30e50, __ptr=0x0000000005404b00) const at unique_ptr.h:85:2
frame #2: 0x00007fffee39edca libonnxruntime.so.1.16.0`std::unique_ptr<onnxruntime::Environment, std::default_delete<onnxruntime::Environment>>::~unique_ptr(this=0x5404b00) at unique_ptr.h:361:17
frame #3: 0x00007fffee39e2ab libonnxruntime.so.1.16.0`OrtEnv::~OrtEnv(this=0x00007ffff4c30e50) at ort_env.cc:43:1
frame #4: 0x00007fffee39fa96 libonnxruntime.so.1.16.0`std::default_delete<OrtEnv>::operator()(this=0x00007fffefff8f78, __ptr=0x00007ffff4c30e50) const at unique_ptr.h:85:2
frame #5: 0x00007fffee39f394 libonnxruntime.so.1.16.0`std::unique_ptr<OrtEnv, std::default_delete<OrtEnv>>::~unique_ptr(this=0x7ffff4c30e50) at unique_ptr.h:361:17
frame #6: 0x00007ffff78574b5 libc.so.6`__run_exit_handlers + 261
frame #7: 0x00007ffff7857630 libc.so.6`exit + 32
frame #8: 0x00007ffff783feb7 libc.so.6`__libc_start_call_main + 135
frame #9: 0x00007ffff783ff60 libc.so.6`__libc_start_main@@GLIBC_2.34 + 128
frame #10: 0x0000000000abbdee node`_start + 46
```
After this change, OrtEnv will be deleted before the main function
returns and nodejs is still alive.
preetha-intel
pushed a commit
that referenced
this pull request
Jul 29, 2024
### Description Security fuzz test with address sanitizer found several bugs
hdharpure9922
pushed a commit
that referenced
this pull request
Jul 28, 2026
…rosoft#29880) ### Description `import onnxruntime` segfaults during `dlopen` of `onnxruntime_pybind11_state.so` on Linux. This removes the global initializer in `onnxruntime_pybind_state.cc` that eagerly calls `Env::Default()`, and resolves the platform `Env` on first use at its two call sites instead. ### Motivation and Context The module has a namespace-scope dynamic initializer: ```cpp static Env& platform_env = Env::Default(); ``` Since POSIX telemetry landed (microsoft#27379), `Env::Default()` constructs `PosixEnv`, whose `PosixTelemetry` member initializes the 1DS SDK in its constructor. That path reads `defaultRuntimeConfig`, a namespace-scope `static ILogConfiguration` defined in the 1DS SDK's `RuntimeConfig_Default.hpp`, which lives in a **different translation unit of the same shared library**. Dynamic initialization order across translation units is unspecified, and the pybind TU's initializer runs first. `Variant::merge_map` therefore iterates a still zero-initialized `std::map`: `_M_node_count == 0`, but `_M_header._M_left` is `nullptr` rather than self-pointing, so `begin() != end()` and the loop dereferences null. Textbook static initialization order fiasco. Backtrace (Release build relinked without `--strip-all` to recover symbols): ``` #0 std::map<..., Variant>::lower_bound stl_map.h:1307 #1 std::map<..., Variant>::operator[] stl_map.h:509 #2 Variant::merge_map VariantType.hpp:508 #3 RuntimeConfig_Default::RuntimeConfig_Default RuntimeConfig_Default.hpp:97 #4 LogManagerImpl::LogManagerImpl LogManagerImpl.cpp:183 #5 LogManagerFactory::Create LogManagerFactory.cpp:36 #6 LogManagerFactory::lease #7 LogManagerFactory::Get LogManagerFactory.hpp:71 #8 LogManagerProvider::Get LogManagerProvider.cpp:16 #9 onnxruntime::PosixTelemetry::Initialize() #10 onnxruntime::PosixTelemetry::PosixTelemetry() #11 onnxruntime::(anonymous namespace)::PosixEnv::PosixEnv() #12 onnxruntime::Env::Default() #13 _GLOBAL__sub_I_onnxruntime_pybind_state.cc #14 call_init elf/dl-init.c:74 ... #21 _dl_open elf/dl-open.c:905 ``` The crash is independent of `LD_LIBRARY_PATH` and `CUDA_VISIBLE_DEVICES` — it happens before any ORT runtime code runs. `libonnxruntime.so` is unaffected because it contains no global initializer that reaches `Env::Default()`, which is why C/C++ and onnxruntime-genai consumers do not see it. Both remaining uses of `platform_env` are inside pybind lambdas that run long after load, so calling `Env::Default()` there is safe. This also removes the now-stale `TODO: we may delay-init this variable` and a pre-existing `#pragma warning(push)` that should have been `pop`. Note for follow-up: `Env::Default()` is now unsafe to call from any dynamic initializer. This was the only such call site in the tree, but hardening `PosixTelemetry` to defer SDK initialization out of its constructor would remove the hazard entirely. ### Tests Verified on a Linux CUDA 13 Release build (`onnxruntime_USE_TELEMETRY` enabled): - `dlopen` of `onnxruntime_pybind11_state.so` succeeds (previously SIGSEGV). - `import onnxruntime` reports the version and `['CUDAExecutionProvider', 'CPUExecutionProvider']`. - `onnxruntime.enable_telemetry_events()` / `disable_telemetry_events()` — the two call sites changed here — work. - CPU and CUDA inference sessions produce correct results. - Reproduced and verified with `LD_LIBRARY_PATH` unset and `CUDA_VISIBLE_DEVICES` empty.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description: Describe your changes.
Motivation and Context