Skip to content

Make GGML_SYCL_F16=ON the default - #23996

Merged
ggerganov merged 5 commits into
ggml-org:masterfrom
aicss-genai:sycl-f16-default
Jun 16, 2026
Merged

Make GGML_SYCL_F16=ON the default#23996
ggerganov merged 5 commits into
ggml-org:masterfrom
aicss-genai:sycl-f16-default

Conversation

@malsbat

@malsbat malsbat commented Jun 1, 2026

Copy link
Copy Markdown
Contributor

Overview

The current default of GGML_SYCL_F16 is OFF. There are significant performance gains to be had during prompt processing by setting it ON: improvement on models measured averaged 2.43x (ranged from 1.48x to 3.41x) for prompt processing. Token generation measured no improvement (average of 1.00x, ranged from 0.97x to 1.02x).

At the same time, llama-perplexity showed no meaningful degradation on the wiki.test.raw, winogrande, or hellaswag benchmarks.

Additional information

Measurements below were captured on B70.

Performance measurements

Metric Value
Average prefill speedup (F16/F32) 2.43x
Average decode speedup (F16/F32) 1.00x
Model Task Tokens F32 (tok/s) F16 (tok/s) Speedup
DeepSeek-R1-Qwen-32B-Q4 pp 512 255.34 ±0.00 702.80 ±0.00 2.75x
DeepSeek-R1-Qwen-32B-Q4 pp 1024 269.73 ±0.00 850.26 ±0.00 3.15x
DeepSeek-R1-Qwen-32B-Q4 pp 2048 261.36 ±0.00 816.25 ±0.00 3.12x
DeepSeek-R1-Qwen-32B-Q4 pp 4096 237.66 ±0.00 625.70 ±0.00 2.63x
DeepSeek-R1-Qwen-32B-Q4 pp 8192 201.03 ±0.00 428.58 ±0.00 2.13x
DeepSeek-R1-Qwen-32B-Q4 tg 128 20.99 ±0.00 20.97 ±0.00 1.00x
DeepSeek-R1-Qwen-32B-Q4 tg 256 20.91 ±0.00 20.93 ±0.00 1.00x
DeepSeek-R1-Qwen-32B-Q4 tg 512 20.76 ±0.00 20.81 ±0.00 1.00x
DeepSeek-R1-Qwen-32B-Q4 tg 1024 20.52 ±0.00 20.48 ±0.00 1.00x
Gemma-2-9B pp 512 861.32 ±0.00 1993.83 ±0.00 2.31x
Gemma-2-9B pp 1024 884.10 ±0.00 2257.81 ±0.00 2.55x
Gemma-2-9B pp 2048 842.96 ±0.00 1980.15 ±0.00 2.35x
Gemma-2-9B pp 4096 720.86 ±0.00 1423.74 ±0.00 1.98x
Gemma-2-9B pp 8192 566.30 ±0.00 929.83 ±0.00 1.64x
Gemma-2-9B tg 128 58.58 ±0.00 58.79 ±0.00 1.00x
Gemma-2-9B tg 256 58.35 ±0.00 58.46 ±0.00 1.00x
Gemma-2-9B tg 512 58.10 ±0.00 58.19 ±0.00 1.00x
Gemma-2-9B tg 1024 55.93 ±0.00 56.17 ±0.00 1.00x
Llama-3.1-8B-Q8 pp 512 1045.67 ±0.00 2745.04 ±0.00 2.63x
Llama-3.1-8B-Q8 pp 1024 1103.75 ±0.00 3091.26 ±0.00 2.80x
Llama-3.1-8B-Q8 pp 2048 1034.44 ±0.00 2621.40 ±0.00 2.53x
Llama-3.1-8B-Q8 pp 4096 890.18 ±0.00 1903.68 ±0.00 2.14x
Llama-3.1-8B-Q8 pp 8192 695.90 ±0.00 1209.49 ±0.00 1.74x
Llama-3.1-8B-Q8 tg 128 56.27 ±0.00 56.22 ±0.00 1.00x
Llama-3.1-8B-Q8 tg 256 56.30 ±0.00 56.23 ±0.00 1.00x
Llama-3.1-8B-Q8 tg 512 56.18 ±0.00 56.10 ±0.00 1.00x
Llama-3.1-8B-Q8 tg 1024 55.59 ±0.00 55.50 ±0.00 1.00x
Llama-3.2-3B pp 512 2485.88 ±0.00 4973.16 ±0.00 2.00x
Llama-3.2-3B pp 1024 2455.36 ±0.00 5407.97 ±0.00 2.20x
Llama-3.2-3B pp 2048 2196.37 ±0.00 4455.35 ±0.00 2.03x
Llama-3.2-3B pp 4096 1793.81 ±0.00 3118.76 ±0.00 1.74x
Llama-3.2-3B pp 8192 1305.02 ±0.00 1933.71 ±0.00 1.48x
Llama-3.2-3B tg 128 144.29 ±0.00 145.12 ±0.00 1.01x
Llama-3.2-3B tg 256 144.03 ±0.00 144.28 ±0.00 1.00x
Llama-3.2-3B tg 512 142.76 ±0.00 142.95 ±0.00 1.00x
Llama-3.2-3B tg 1024 137.64 ±0.00 138.03 ±0.00 1.00x
Mistral-Nemo-12B pp 512 702.92 ±0.00 1786.92 ±0.00 2.54x
Mistral-Nemo-12B pp 1024 740.79 ±0.00 2064.76 ±0.00 2.79x
Mistral-Nemo-12B pp 2048 705.70 ±0.00 1890.57 ±0.00 2.68x
Mistral-Nemo-12B pp 4096 621.40 ±0.00 1396.77 ±0.00 2.25x
Mistral-Nemo-12B pp 8192 501.28 ±0.00 915.55 ±0.00 1.83x
Mistral-Nemo-12B tg 128 54.60 ±0.00 54.31 ±0.00 0.99x
Mistral-Nemo-12B tg 256 54.46 ±0.00 53.95 ±0.00 0.99x
Mistral-Nemo-12B tg 512 54.34 ±0.00 53.74 ±0.00 0.99x
Mistral-Nemo-12B tg 1024 53.69 ±0.00 53.09 ±0.00 0.99x
Mistral-Small-24B pp 512 360.64 ±0.00 1078.51 ±0.00 2.99x
Mistral-Small-24B pp 1024 381.18 ±0.00 1308.88 ±0.00 3.43x
Mistral-Small-24B pp 2048 377.68 ±0.00 1303.19 ±0.00 3.45x
Mistral-Small-24B pp 4096 351.74 ±0.00 1045.46 ±0.00 2.97x
Mistral-Small-24B pp 8192 309.36 ±0.00 753.30 ±0.00 2.44x
Mistral-Small-24B tg 128 29.70 ±0.00 30.02 ±0.00 1.01x
Mistral-Small-24B tg 256 29.63 ±0.00 29.92 ±0.00 1.01x
Mistral-Small-24B tg 512 29.50 ±0.00 29.85 ±0.00 1.01x
Mistral-Small-24B tg 1024 29.30 ±0.00 29.63 ±0.00 1.01x
Phi-3.5-mini-3.8B pp 512 1853.10 ±0.00 4075.13 ±0.00 2.20x
Phi-3.5-mini-3.8B pp 1024 1843.66 ±0.00 4235.75 ±0.00 2.30x
Phi-3.5-mini-3.8B pp 2048 1693.64 ±0.00 3417.81 ±0.00 2.02x
Phi-3.5-mini-3.8B pp 4096 1367.47 ±0.00 2397.30 ±0.00 1.75x
Phi-3.5-mini-3.8B pp 8192 992.83 ±0.00 1471.76 ±0.00 1.48x
Phi-3.5-mini-3.8B tg 128 134.95 ±0.00 135.52 ±0.00 1.00x
Phi-3.5-mini-3.8B tg 256 134.57 ±0.00 134.70 ±0.00 1.00x
Phi-3.5-mini-3.8B tg 512 130.86 ±0.00 131.12 ±0.00 1.00x
Phi-3.5-mini-3.8B tg 1024 124.89 ±0.00 125.02 ±0.00 1.00x
Qwen2.5-14B-Q4 pp 512 565.78 ±0.00 1393.04 ±0.00 2.46x
Qwen2.5-14B-Q4 pp 1024 595.27 ±0.00 1595.79 ±0.00 2.68x
Qwen2.5-14B-Q4 pp 2048 560.66 ±0.00 1433.56 ±0.00 2.56x
Qwen2.5-14B-Q4 pp 4096 482.25 ±0.00 1023.94 ±0.00 2.12x
Qwen2.5-14B-Q4 pp 8192 378.14 ±0.00 655.40 ±0.00 1.73x
Qwen2.5-14B-Q4 tg 128 43.64 ±0.00 43.55 ±0.00 1.00x
Qwen2.5-14B-Q4 tg 256 43.43 ±0.00 43.37 ±0.00 1.00x
Qwen2.5-14B-Q4 tg 512 43.16 ±0.00 43.07 ±0.00 1.00x
Qwen2.5-14B-Q4 tg 1024 42.25 ±0.00 42.19 ±0.00 1.00x
Qwen2.5-14B-Q8 pp 512 557.42 ±0.00 1461.46 ±0.00 2.62x
Qwen2.5-14B-Q8 pp 1024 590.94 ±0.00 1645.93 ±0.00 2.79x
Qwen2.5-14B-Q8 pp 2048 559.19 ±0.00 1432.49 ±0.00 2.56x
Qwen2.5-14B-Q8 pp 4096 481.75 ±0.00 1018.36 ±0.00 2.11x
Qwen2.5-14B-Q8 pp 8192 377.23 ±0.00 647.09 ±0.00 1.72x
Qwen2.5-14B-Q8 tg 128 28.59 ±0.00 28.61 ±0.00 1.00x
Qwen2.5-14B-Q8 tg 256 28.57 ±0.00 28.61 ±0.00 1.00x
Qwen2.5-14B-Q8 tg 512 28.43 ±0.00 28.46 ±0.00 1.00x
Qwen2.5-14B-Q8 tg 1024 28.04 ±0.00 28.06 ±0.00 1.00x
Qwen2.5-32B-Q4 pp 512 250.16 ±0.00 703.77 ±0.00 2.81x
Qwen2.5-32B-Q4 pp 1024 262.57 ±0.00 848.11 ±0.00 3.23x
Qwen2.5-32B-Q4 pp 2048 255.08 ±0.00 810.17 ±0.00 3.18x
Qwen2.5-32B-Q4 pp 4096 232.17 ±0.00 624.48 ±0.00 2.69x
Qwen2.5-32B-Q4 pp 8192 196.77 ±0.00 428.02 ±0.00 2.18x
Qwen2.5-32B-Q4 tg 128 20.58 ±0.00 20.94 ±0.00 1.02x
Qwen2.5-32B-Q4 tg 256 20.53 ±0.00 20.86 ±0.00 1.02x
Qwen2.5-32B-Q4 tg 512 20.43 ±0.00 20.73 ±0.00 1.01x
Qwen2.5-32B-Q4 tg 1024 20.15 ±0.00 20.46 ±0.00 1.02x
Qwen2.5-32B-Q6 pp 512 252.12 ±0.00 756.21 ±0.00 3.00x
Qwen2.5-32B-Q6 pp 1024 267.10 ±0.00 858.42 ±0.00 3.21x
Qwen2.5-32B-Q6 pp 2048 258.91 ±0.00 809.55 ±0.00 3.13x
Qwen2.5-32B-Q6 pp 4096 236.23 ±0.00 619.63 ±0.00 2.62x
Qwen2.5-32B-Q6 pp 8192 200.27 ±0.00 423.10 ±0.00 2.11x
Qwen2.5-32B-Q6 tg 128 12.83 ±0.00 12.49 ±0.00 0.97x
Qwen2.5-32B-Q6 tg 256 12.76 ±0.00 12.44 ±0.00 0.97x
Qwen2.5-32B-Q6 tg 512 12.76 ±0.00 12.41 ±0.00 0.97x
Qwen2.5-32B-Q6 tg 1024 12.66 ±0.00 12.32 ±0.00 0.97x
Qwen2.5-7B pp 512 1142.70 ±0.00 2918.44 ±0.00 2.55x
Qwen2.5-7B pp 1024 1213.82 ±0.00 3246.44 ±0.00 2.67x
Qwen2.5-7B pp 2048 1165.09 ±0.00 3090.48 ±0.00 2.65x
Qwen2.5-7B pp 4096 1021.99 ±0.00 2300.18 ±0.00 2.25x
Qwen2.5-7B pp 8192 821.50 ±0.00 1505.63 ±0.00 1.83x
Qwen2.5-7B tg 128 82.13 ±0.00 82.39 ±0.00 1.00x
Qwen2.5-7B tg 256 81.62 ±0.00 82.00 ±0.00 1.00x
Qwen2.5-7B tg 512 80.06 ±0.00 80.47 ±0.00 1.01x
Qwen2.5-7B tg 1024 78.68 ±0.00 79.11 ±0.00 1.01x
Qwen3-8B pp 512 1040.88 ±0.00 2483.45 ±0.00 2.39x
Qwen3-8B pp 1024 1060.96 ±0.00 2812.53 ±0.00 2.65x
Qwen3-8B pp 2048 1010.09 ±0.00 2409.60 ±0.00 2.39x
Qwen3-8B pp 4096 857.38 ±0.00 1722.44 ±0.00 2.01x
Qwen3-8B pp 8192 658.54 ±0.00 1095.43 ±0.00 1.66x
Qwen3-8B tg 128 76.29 ±0.00 76.65 ±0.00 1.00x
Qwen3-8B tg 256 76.07 ±0.00 76.53 ±0.00 1.01x
Qwen3-8B tg 512 75.89 ±0.00 76.19 ±0.00 1.00x
Qwen3-8B tg 1024 74.66 ±0.00 75.13 ±0.00 1.01x
Qwen3.5-9B-Q4 pp 512 1017.20 ±0.00 2390.18 ±0.00 2.35x
Qwen3.5-9B-Q4 pp 1024 1047.11 ±0.00 2727.16 ±0.00 2.60x
Qwen3.5-9B-Q4 pp 2048 1066.86 ±0.00 2749.60 ±0.00 2.58x
Qwen3.5-9B-Q4 pp 4096 1020.11 ±0.00 2504.86 ±0.00 2.46x
Qwen3.5-9B-Q4 pp 8192 943.54 ±0.00 2112.99 ±0.00 2.24x
Qwen3.5-9B-Q4 tg 128 65.74 ±0.00 65.56 ±0.00 1.00x
Qwen3.5-9B-Q4 tg 256 65.47 ±0.00 65.27 ±0.00 1.00x
Qwen3.5-9B-Q4 tg 512 65.42 ±0.00 65.15 ±0.00 1.00x
Qwen3.5-9B-Q4 tg 1024 65.30 ±0.00 65.08 ±0.00 1.00x
image

Perplexity measurements

hellaswag

Model Metric F32 Score F16 Score Delta (F16-F32) Ratio (F16/F32)
DeepSeek-R1-Qwen-32B-Q4 hellaswag 80.0000 80.0000 +0.0000 1.0000
Gemma-2-9B hellaswag 78.5000 78.5000 +0.0000 1.0000
Llama-3.1-8B-Q8 hellaswag 78.5000 78.5000 +0.0000 1.0000
Llama-3.2-3B hellaswag 73.0000 73.0000 +0.0000 1.0000
Mistral-Nemo-12B hellaswag 81.5000 81.5000 +0.0000 1.0000
Mistral-Small-24B hellaswag 82.0000 82.0000 +0.0000 1.0000
Phi-3.5-mini-3.8B hellaswag 76.5000 76.7500 +0.2500 1.0033
Qwen2.5-14B-Q4 hellaswag 82.7500 82.7500 +0.0000 1.0000
Qwen2.5-14B-Q8 hellaswag 82.7500 82.7500 +0.0000 1.0000
Qwen2.5-32B-Q4 hellaswag 84.7500 84.7500 +0.0000 1.0000
Qwen2.5-32B-Q6 hellaswag 84.7500 84.7500 +0.0000 1.0000
Qwen2.5-7B hellaswag 80.0000 80.0000 +0.0000 1.0000
Qwen3-8B hellaswag 73.7500 73.7500 +0.0000 1.0000
Qwen3.5-9B-Q4 hellaswag 78.0000 78.0000 +0.0000 1.0000
image

wiki.test.raw

Model Metric F32 PPL F16 PPL Delta (F16-F32) Ratio (F16/F32)
DeepSeek-R1-Qwen-32B-Q4 perplexity 7.1700 7.1696 -0.0004 0.9999
Gemma-2-9B perplexity 8.7995 8.7993 -0.0002 1.0000
Llama-3.1-8B-Q8 perplexity 7.3245 7.3245 +0.0000 1.0000
Llama-3.2-3B perplexity 10.7252 10.7252 +0.0000 1.0000
Mistral-Nemo-12B perplexity 6.4357 6.4357 +0.0000 1.0000
Mistral-Small-24B perplexity 5.9889 5.9888 -0.0001 1.0000
Phi-3.5-mini-3.8B perplexity 6.6916 6.6916 +0.0000 1.0000
Qwen2.5-14B-Q4 perplexity 6.1145 6.1145 +0.0000 1.0000
Qwen2.5-14B-Q8 perplexity 5.9765 5.9766 +0.0001 1.0000
Qwen2.5-32B-Q4 perplexity 5.6140 5.6142 +0.0002 1.0000
Qwen2.5-32B-Q6 perplexity 5.5257 5.5257 +0.0000 1.0000
Qwen2.5-7B perplexity 7.9973 7.9974 +0.0001 1.0000
Qwen3-8B perplexity 10.4663 10.4663 +0.0000 1.0000
Qwen3.5-9B-Q4 perplexity 8.3137 8.3137 +0.0000 1.0000
image

winogrande

Model Metric F32 Score F16 Score Delta (F16-F32) Ratio (F16/F32)
DeepSeek-R1-Qwen-32B-Q4 winogrande 75.5328 75.5328 +0.0000 1.0000
Gemma-2-9B winogrande 76.4799 76.4009 -0.0790 0.9990
Llama-3.1-8B-Q8 winogrande 73.5596 73.5596 +0.0000 1.0000
Llama-3.2-3B winogrande 68.1926 68.1926 +0.0000 1.0000
Mistral-Nemo-12B winogrande 76.9534 76.9534 +0.0000 1.0000
Mistral-Small-24B winogrande 79.2423 79.2423 +0.0000 1.0000
Phi-3.5-mini-3.8B winogrande 73.7964 73.7964 +0.0000 1.0000
Qwen2.5-14B-Q4 winogrande 74.1910 74.0331 -0.1579 0.9979
Qwen2.5-14B-Q8 winogrande 73.8753 73.8753 +0.0000 1.0000
Qwen2.5-32B-Q4 winogrande 72.5335 72.4546 -0.0789 0.9989
Qwen2.5-32B-Q6 winogrande 72.9282 72.8493 -0.0789 0.9989
Qwen2.5-7B winogrande 71.0339 71.0339 +0.0000 1.0000
Qwen3-8B winogrande 68.5083 68.4294 -0.0789 0.9988
Qwen3.5-9B-Q4 winogrande 71.7443 71.7443 +0.0000 1.0000
image

Requirements

malsbat added 2 commits May 13, 2026 23:32
Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>
Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>
@malsbat
malsbat requested a review from ngxson as a code owner June 1, 2026 21:23
@github-actions github-actions Bot added documentation Improvements or additions to documentation examples devops improvements to build systems and github actions ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language labels Jun 1, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Jun 1, 2026

Copy link
Copy Markdown

Hi @malsbat, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

@arthw arthw left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SYCL support fp32 and fp16 for two reasons:

  1. fp32 will provide better accuracy than fp16.
  2. fp32 show better performance than fp16 in some cases.

I agree to recommend to use fp16 as default.

But not change the default value of parameter GGML_SYCL_F16 and the default values in script: That will impact existed CI in all user development.

@malsbat

malsbat commented Jun 2, 2026

Copy link
Copy Markdown
Contributor Author

@arthw, thanks for reviewing. My measurements are quite different than what you're seeing - I see better performance with fp16 and no meaningful degradation in quality. Can you tell me which models and devices you're seeing better results with fp32? I'd like to add these to my local benchmarks if possible.

@arthw

arthw commented Jun 3, 2026

Copy link
Copy Markdown
Contributor

@malsbat
We collect the performance data and saved in #23313.

We can see the PP of FP16 is about double performance of FP32, but the TG is similar (FP32 is little bigger than FP16 in more cases).

I think FP16 has half bandwidth than FP32 in PP, for TG FP32 code path use fp16 data type in computing in fact.

I'm considering to merge FP16 and FP32 code path to one in the future. Like PP is using FP16 and TG is using FP32.

Anyway, I think recommend FP16 is good idea.
We only need to change the recommended mode in SYCL.md.
To avoid to impact the existed CI/docker/delivery, I suggest not changing the default value and script for it.
The dockerfile, it's OK to set FP16=on.

Thank you!

F16 remains explictly set for example and Dockerfile builds.

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>
@malsbat

malsbat commented Jun 3, 2026

Copy link
Copy Markdown
Contributor Author

Thanks @arthw, I may have misunderstood your earlier comment. I've updated the PR to leave F32 the default in CMakeLists.txt and explicitly set F16=ON in the other locations.

I did notice that https://github.com/ggml-org/llama.cpp/blob/master/ci/run.sh#L113 explicitly sets F16=ON, so I'm unsure of your comment about impacting existing CI. Please let me know if this PR matches what you intended.

The main change I want is to the Dockerfile so that end-users get the best performance when using the prebuilt images.

@arthw

arthw commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

@malsbat
I see your update.

ci/run.sh - In fact, SYCL backend only CI is only building, not execute this script.
I will execute the CI locally, but not using this script.

build.sh and win-build-sycl.bat: I hope keep the default behavior (fp32), so that the existed CI (different users locally) won't be impacted.

I'm considering to merge fp32 and fp16 code paths in different function/OPs, to get the best performance.
The flag fp32 and fp16 will disappear in code in the future.

So we won't need to modify more on this issue now.

Thank you!

@malsbat

malsbat commented Jun 8, 2026

Copy link
Copy Markdown
Contributor Author

Should I close this PR then?

@arthw

arthw commented Jun 9, 2026

Copy link
Copy Markdown
Contributor

No! :)
Only rollback: build.sh and win-build-sycl.bat.

Thank you!

@malsbat

malsbat commented Jun 9, 2026

Copy link
Copy Markdown
Contributor Author

Thanks for clarifying, I will update this PR shortly.

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>
@malsbat

malsbat commented Jun 9, 2026

Copy link
Copy Markdown
Contributor Author

Updated, please review when you can

@arthw arthw left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's good job!

Thank you!

@arthw arthw added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Jun 11, 2026
@ggerganov

Copy link
Copy Markdown
Member

@arthw Are you sure you want to make this change - F16 accumulators will definitely overflow in some cases, so generally not recommended.

@arthw

arthw commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

@arthw Are you sure you want to make this change - F16 accumulators will definitely overflow in some cases, so generally not recommended.

@ggerganov
The overflow in FP16 is present on some cases.
But there are good results of FP16 in many cases. Please refer to #23313.

I think FP16 will be chosen by most user if there is no such overflow issue.
We didn't see the abnormal output in FP16 case for a long time.

This PR is used to change to FP16 in docker building and recommend FP16 in guide.
User can switch to FP32 in any time.

In same time, the docker user will feedback the overflow issue as docker use FP16 as default.
It will help us to collect the data of FP16.
The impact will be limited.

Another issue, I'm considering to merge FP32 and FP16 in SYCL backend: choose the performance better part in kernel.
But as your mentioned, there is risk to remove FP32 in some OPs.
I will think about it.

Thank you!

@ggerganov
ggerganov merged commit 4196b47 into ggml-org:master Jun 16, 2026
3 checks passed
papamoose pushed a commit to papamoose/llama.cpp that referenced this pull request Jun 27, 2026
* Add -cl-fp32-correctly-rounded-divide-sqrt to F16=ON builds

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Make GGML_SYCL_F16=ON the default

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Leave F32 the default

F16 remains explictly set for example and Dockerfile builds.

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Revert changes to examples/sycl/build scripts

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

---------

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>
adrianhoehne pushed a commit to adrianhoehne/llama.cpp that referenced this pull request Jul 5, 2026
* Add -cl-fp32-correctly-rounded-divide-sqrt to F16=ON builds

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Make GGML_SYCL_F16=ON the default

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Leave F32 the default

F16 remains explictly set for example and Dockerfile builds.

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Revert changes to examples/sycl/build scripts

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

---------

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

devops improvements to build systems and github actions documentation Improvements or additions to documentation examples ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants