Skip to content

[WebNN EP] Automatically use ml-tensor for outputs - #24282

Merged
fs-eire merged 5 commits into
microsoft:mainfrom
egalli:promote_outputs
Apr 16, 2025
Merged

[WebNN EP] Automatically use ml-tensor for outputs#24282
fs-eire merged 5 commits into
microsoft:mainfrom
egalli:promote_outputs

Conversation

@egalli

@egalli egalli commented Apr 2, 2025

Copy link
Copy Markdown
Contributor

Description

If it would improve performance, this patch moves outputs to MLTensor backed Tensors.

Motivation and Context

We are currently performing an extra copy on output tensors located in the CPU when using the WebNN EP (MLTensor -(copy)-> wasm heap -(copy)-> JS). This patch removes this copy by moving the readback to JS instead of wasm. As an extra benefit, we can also start the readbacks and wait for them in parallel.

This change is similar to #23073

### Description
If it would improve performance, this patch moves outputs to MLTensor backed Tensors.

### Motivation and Context
We are currently performing an extra copy on output tensors located in the CPU when using the WebNN EP (MLTensor -(copy)-> wasm heap -(copy)-> JS). This patch removes this copy by moving the readback to JS instead of wasm. As an extra benefit, we can also start and wait for the readbacks in parallel.
@snnn snnn closed this Apr 3, 2025
@snnn snnn reopened this Apr 3, 2025
@guschmue guschmue added the ep:WebNN WebNN execution provider label Apr 3, 2025
@snnn

snnn commented Apr 3, 2025

Copy link
Copy Markdown
Contributor

/azp run all

@azure-pipelines

Copy link
Copy Markdown
No pipelines are associated with this pull request.

@snnn

snnn commented Apr 3, 2025

Copy link
Copy Markdown
Contributor

/azp run Linux QNN CI Pipeline, Win_TRT_Minimal_CUDA_Test_CI, Windows ARM64 QNN CI Pipeline,Windows GPU Doc Gen CI Pipeline,Windows x64 QNN CI Pipeline

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 5 pipeline(s).

Comment thread js/web/lib/wasm/wasm-core-impl.ts
@fs-eire

fs-eire commented Apr 13, 2025

Copy link
Copy Markdown
Contributor

/azp run Windows ARM64 QNN CI Pipeline,Windows x64 QNN CI Pipeline,Windows GPU Doc Gen CI Pipeline,ONNX Runtime Web CI Pipeline,Win_TRT_Minimal_CUDA_Test_CI,Linux CPU CI Pipeline,Linux CPU Minimal Build E2E CI Pipeline,Linux GPU CI Pipeline,Linux GPU TensorRT CI Pipeline,Linux OpenVINO CI Pipeline

@fs-eire

fs-eire commented Apr 13, 2025

Copy link
Copy Markdown
Contributor

/azp run Linux QNN CI Pipeline,onnxruntime-binary-size-checks-ci-pipeline,Big Models,Linux Android Emulator QNN CI Pipeline,Android CI Pipeline,iOS CI Pipeline,ONNX Runtime React Native CI Pipeline,Linux DNNL CI Pipeline,Linux MIGraphX CI Pipeline,Linux ROCm CI Pipeline

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 5 pipeline(s).

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 7 pipeline(s).

@fs-eire
fs-eire merged commit 1c2225e into microsoft:main Apr 16, 2025
ashrit-ms pushed a commit that referenced this pull request Apr 24, 2025
### Description
If it would improve performance, this patch moves outputs to MLTensor
backed Tensors.

### Motivation and Context
We are currently performing an extra copy on output tensors located in
the CPU when using the WebNN EP (MLTensor -(copy)-> wasm heap -(copy)->
JS). This patch removes this copy by moving the readback to JS instead
of wasm. As an extra benefit, we can also start the readbacks and wait
for them in parallel.

This change is similar to #23073
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ep:WebNN WebNN execution provider

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants