Skip to content

Streaming conversion with no torch#176

Closed
diimdeep wants to merge 6 commits into
ggml-org:masterfrom
diimdeep:streaming
Closed

Streaming conversion with no torch#176
diimdeep wants to merge 6 commits into
ggml-org:masterfrom
diimdeep:streaming

Conversation

@diimdeep

Copy link
Copy Markdown

Drop torch, do not load whole file into memory, process files in parallel and use separate threads for r/w

Comment thread convert-pth-to-ggml.py
q = queue.Queue(maxsize=2)

def writer():
while True:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

while True? Does this function ever return? I don't know if the function exists but maybe something like while !q.atEnd()

Please correct me if I'm wrong. I haven't worked with Python since a year or so.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed!

@gjmulder gjmulder added enhancement New feature or request performance Speed related topics labels Mar 15, 2023
@sw

sw commented Mar 19, 2023

Copy link
Copy Markdown
Contributor

The python dependencies in .devops/full.Dockerfile should also be updated, will conflict with my PR #293.

@ggerganov

Copy link
Copy Markdown
Member

This looks like a very useful addition. Lets give it a priority and merge after resolving the conflicts

@tim-gromeyer

Copy link
Copy Markdown

@ggerganov Any update on this? Because I really do not want to install pytorch on my system (because of memory).

@Green-Sky Green-Sky added the high priority Very important issue label Mar 26, 2023
@sw sw mentioned this pull request Mar 27, 2023
16 tasks
@ggerganov

Copy link
Copy Markdown
Member

This is probably too outdated so closing for now

@ggerganov ggerganov closed this Mar 30, 2023
phuongncn pushed a commit to phuongncn/llama.cpp-gx10-dgx-sparks-deepseekv4 that referenced this pull request Apr 28, 2026
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
drizzt pushed a commit to drizzt/llama.cpp that referenced this pull request Jun 15, 2026
…ILE FA routing (ggml-org#176)

* HIP: fix turbo KV decode crash under graph capture; batch-aware VEC/TILE FA routing

Route small-batch (decode) quantized-KV flash attention through the graph-safe VEC kernel and let large prefill batches fall through to the fast TILE/MMA kernel. Make the f16 dequant temp allocation capture-aware: allocate from the ggml pool while a stream is capturing (no cudaMalloc/cudaFree/cudaStreamSynchronize), keep raw alloc for large eager prefill so the multi-GB buffer is released immediately (gfx1201 has no VMM, the legacy pool would retain it).

Fixes 'FLASH_ATTN_EXT failed: operation not permitted when stream is capturing' with GGML_HIP_GRAPHS=ON and turbo KV types on RDNA4. Tested on gfx1201 (Radeon AI PRO R9700, Windows, HIP SDK 7.1): pp2048 735 t/s (vs 188 t/s without graphs), tg128 22.9 t/s, no decode crash. Possibly related: ggml-org#12.

* fattn (HIP): note pool-retention tradeoff for non-VEC captured decode

Address review on ggml-org#176: document that head_dim==192 / K-stride-mismatch
configs fall through to the TILE/MMA path under capture and pool-alloc the
full f16 dequant buffer, which the legacy pool retains permanently -- a VRAM
tradeoff, not a crash. VEC-eligible head dims (Gemma) never hit this.

---------

Co-authored-by: KaiAtAdesso <KaiAtAdesso@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request high priority Very important issue performance Speed related topics

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants