Streaming conversion with no torch#176
Closed
diimdeep wants to merge 6 commits into
Closed
Conversation
| q = queue.Queue(maxsize=2) | ||
|
|
||
| def writer(): | ||
| while True: |
There was a problem hiding this comment.
while True? Does this function ever return? I don't know if the function exists but maybe something like while !q.atEnd()
Please correct me if I'm wrong. I haven't worked with Python since a year or so.
Contributor
|
The python dependencies in .devops/full.Dockerfile should also be updated, will conflict with my PR #293. |
Member
|
This looks like a very useful addition. Lets give it a priority and merge after resolving the conflicts |
|
@ggerganov Any update on this? Because I really do not want to install pytorch on my system (because of memory). |
Member
|
This is probably too outdated so closing for now |
phuongncn
pushed a commit
to phuongncn/llama.cpp-gx10-dgx-sparks-deepseekv4
that referenced
this pull request
Apr 28, 2026
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
drizzt
pushed a commit
to drizzt/llama.cpp
that referenced
this pull request
Jun 15, 2026
…ILE FA routing (ggml-org#176) * HIP: fix turbo KV decode crash under graph capture; batch-aware VEC/TILE FA routing Route small-batch (decode) quantized-KV flash attention through the graph-safe VEC kernel and let large prefill batches fall through to the fast TILE/MMA kernel. Make the f16 dequant temp allocation capture-aware: allocate from the ggml pool while a stream is capturing (no cudaMalloc/cudaFree/cudaStreamSynchronize), keep raw alloc for large eager prefill so the multi-GB buffer is released immediately (gfx1201 has no VMM, the legacy pool would retain it). Fixes 'FLASH_ATTN_EXT failed: operation not permitted when stream is capturing' with GGML_HIP_GRAPHS=ON and turbo KV types on RDNA4. Tested on gfx1201 (Radeon AI PRO R9700, Windows, HIP SDK 7.1): pp2048 735 t/s (vs 188 t/s without graphs), tg128 22.9 t/s, no decode crash. Possibly related: ggml-org#12. * fattn (HIP): note pool-retention tradeoff for non-VEC captured decode Address review on ggml-org#176: document that head_dim==192 / K-stride-mismatch configs fall through to the TILE/MMA path under capture and pool-alloc the full f16 dequant buffer, which the legacy pool retains permanently -- a VRAM tradeoff, not a crash. VEC-eligible head dims (Gemma) never hit this. --------- Co-authored-by: KaiAtAdesso <KaiAtAdesso@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Drop torch, do not load whole file into memory, process files in parallel and use separate threads for r/w