llama.cpp

Author	SHA1	Message	Date
shahondin1624	2907ee9830	turboquant: post-merge integration fixes from test validation CI (sycl) / ubuntu-24-sycl (fp16, ON) (push) Has been cancelled Details CI (sycl) / ubuntu-24-sycl (fp32, OFF) (push) Has been cancelled Details CI (sycl) / windows-latest-sycl (push) Has been cancelled Details CI (virtgpu) / ubuntu-24-virtgpu (push) Has been cancelled Details Check vendor / check-vendor (push) Has been cancelled Details CI (vulkan) / ubuntu-24-vulkan-llvmpipe (push) Has been cancelled Details CI (3rd-party) / ubuntu-24-llguidance (push) Has been cancelled Details CI (apple) / macOS-latest-ios (push) Has been cancelled Details CI (apple) / macos-latest-ios-xcode (push) Has been cancelled Details CI (apple) / macOS-latest-tvos (push) Has been cancelled Details CI (apple) / macOS-latest-visionos (push) Has been cancelled Details CI (cann) / openEuler-latest-cann (aarch64, Release, 310p, off) (push) Has been cancelled Details CI (cann) / openEuler-latest-cann (aarch64, Release, 910b, off) (push) Has been cancelled Details CI (cann) / openEuler-latest-cann (aarch64, Release, 910b, on) (push) Has been cancelled Details CI (cann) / openEuler-latest-cann (x86, Release, 310p, off) (push) Has been cancelled Details CI (cann) / openEuler-latest-cann (x86, Release, 910b, off) (push) Has been cancelled Details CI (cann) / openEuler-latest-cann (x86, Release, 910b, on) (push) Has been cancelled Details CI (cross) / debian-13-loongarch64-cpu-cross (push) Has been cancelled Details CI (cross) / debian-13-loongarch64-vulkan-cross (push) Has been cancelled Details CI (cross) / ubuntu-24-riscv64-cpu-spacemit-ime-cross (push) Has been cancelled Details CI (openvino) / ubuntu-24-openvino-CPU (push) Has been cancelled Details CI (openvino) / ubuntu-24-openvino-GPU (push) Has been cancelled Details CI (riscv) / ubuntu-riscv64-native-sanitizer (Debug, ADDRESS) (push) Has been cancelled Details CI (riscv) / ubuntu-riscv64-native-sanitizer (Debug, THREAD) (push) Has been cancelled Details CI (riscv) / ubuntu-riscv64-native-sanitizer (Debug, UNDEFINED) (push) Has been cancelled Details Check Pre-Tokenizer Hashes / pre-tokenizer-hashes (push) Has been cancelled Details flake8 Lint / Lint (push) Has been cancelled Details CI (apple) / macOS-latest-swift (generic/platform=iOS) (push) Has been cancelled Details CI (apple) / macOS-latest-swift (generic/platform=macOS) (push) Has been cancelled Details CI (apple) / macOS-latest-swift (generic/platform=tvOS) (push) Has been cancelled Details Python check requirements.txt / check-requirements (push) Has been cancelled Details Python Type-Check / python type-check (push) Has been cancelled Details CI (snapdragon) / android-ndk-snapdragon (push) Failing after 2m33s Details CI (android) / android (push) Failing after 4m51s Details CI (android) / android-ndk (push) Failing after 4s Details CI (sanitize) / ubuntu-latest-sanitizer (Debug, ADDRESS) (push) Failing after 14s Details CI (sanitize) / ubuntu-latest-sanitizer (Debug, THREAD) (push) Failing after 8s Details CI (sanitize) / ubuntu-latest-sanitizer (Debug, UNDEFINED) (push) Failing after 9s Details CI (UI) / Build static output (push) Failing after 8m40s Details CI (UI) / UI Checks (push) Has been skipped Details CI (UI) / E2E Tests (push) Has been skipped Details CI (snapdragon) / linux-iot-snapdragon (push) Failing after 3m10s Details CI (snapdragon) / Test on QDC Device (QCS9075M) (push) Has been skipped Details CI (snapdragon) / Test on QDC Device (SM8750) (push) Has been skipped Details CI (snapdragon) / Test on QDC Device (SM8850) (push) Has been skipped Details CI / build-cmake-pkg (push) Successful in 15m28s Details CI / android-arm64 (push) Failing after 10s Details CI / ubuntu-latest-rpc (push) Failing after 8s Details CI / ubuntu-latest-cuda (push) Failing after 4m22s Details Release / android-arm64 (push) Failing after 1m10s Details Server (sanitize) / server (RelWithDebInfo, ADDRESS) (push) Failing after 32s Details Server (sanitize) / server (RelWithDebInfo, UNDEFINED) (push) Failing after 4s Details Server / server (default) (push) Failing after 5s Details Server / server (backend-sampling) (push) Failing after 4s Details CI (self-hosted) / ggml-ci-intel-openvino-gpu-low-perf (push) Has been cancelled Details CI (self-hosted) / Determine tag name (push) Has been cancelled Details CI (self-hosted) / ggml-ci-nvidia-cuda (push) Has been cancelled Details CI (self-hosted) / ggml-ci-nvidia-vulkan-cm (push) Has been cancelled Details CI (self-hosted) / ggml-ci-nvidia-vulkan-cm2 (push) Has been cancelled Details CI (self-hosted) / ggml-ci-nvidia-webgpu (push) Has been cancelled Details CI (self-hosted) / ggml-ci-mac-metal (push) Has been cancelled Details CI (self-hosted) / ggml-ci-mac-webgpu (push) Has been cancelled Details CI (self-hosted) / ggml-ci-mac-vulkan (push) Has been cancelled Details CI (self-hosted) / ggml-ci-linux-intel-vulkan (push) Has been cancelled Details CI (self-hosted) / ggml-ci-win-intel-vulkan (push) Has been cancelled Details CI / ggml-ci-arm64-cpu-kleidiai-graviton4 (push) Has been cancelled Details CI / macOS-latest-arm64 (push) Has been cancelled Details CI / macOS-latest-x64 (push) Has been cancelled Details CI / macOS-latest-arm64-webgpu (push) Has been cancelled Details CI / ubuntu-cpu (arm64, ubuntu-24.04-arm) (push) Has been cancelled Details CI / ubuntu-cpu (ppc64le, ubuntu-24.04-ppc64le) (push) Has been cancelled Details CI / ubuntu-24-vulkan (arm64, ubuntu-24.04-arm) (push) Has been cancelled Details CI / ubuntu-24-vulkan (x64, ubuntu-24.04) (push) Has been cancelled Details CI / windows-latest (x64, openblas-x64, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON -DGGML_OPENMP=OFF -DGGML_BLAS=ON -DG… (push) Has been cancelled Details CI / windows-latest (x64, vulkan-x64, -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON -DGGML_VULKAN=ON) (push) Has been cancelled Details CI / windows-2022-cuda (12.4) (push) Has been cancelled Details CI / ubuntu-cpu (s390x, ubuntu-24.04-s390x) (push) Has been cancelled Details CI / ubuntu-cpu (x64, ubuntu-22.04) (push) Has been cancelled Details CI / ubuntu-24-webgpu (push) Has been cancelled Details CI / ubuntu-24-webgpu-wasm (push) Has been cancelled Details CI / ubuntu-22-hip (push) Has been cancelled Details CI / ubuntu-22-musa (push) Has been cancelled Details CI / windows-latest (arm64, llvm-arm64, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/arm64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON) (push) Has been cancelled Details Release / ubuntu-22-rocm (7.2.1, x64, gfx908;gfx90a;gfx942;gfx1030;gfx1100;gfx1101;gfx1102;gfx1151;gfx1150;gfx1200;gfx1201) (push) Has been cancelled Details CI / windows-latest (arm64, llvm-arm64-opencl-adreno, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/arm64-windows-llvm.cmake -DCMAKE_PREFIX_PATH="$env:RUNNER_TEMP/opencl-arm64-release" -DGGML_OPENCL=ON -DGGML_OPENCL_USE_ADRENO_KERNELS=ON) (push) Has been cancelled Details CI / windows-latest (x64, cpu-x64 (static), -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DBUILD_SHARED_LIBS=OFF) (push) Has been cancelled Details CI / windows-latest-hip (push) Has been cancelled Details CI / ubuntu-cpu-riscv64-native (push) Has been cancelled Details CI / ggml-ci-x64-cpu-low-perf (push) Has been cancelled Details CI / ggml-ci-arm64-cpu-low-perf (push) Has been cancelled Details CI / ggml-ci-x64-cpu-high-perf (push) Has been cancelled Details CI / ggml-ci-arm64-cpu-high-perf (push) Has been cancelled Details CI / ggml-ci-arm64-cpu-high-perf-sve (push) Has been cancelled Details CI / ggml-ci-arm64-cpu-kleidiai (push) Has been cancelled Details Code Style Checker / model-naming (push) Has been cancelled Details EditorConfig Checker / editorconfig (push) Has been cancelled Details HIP quality check / ubuntu-22-hip-quality-check (push) Has been cancelled Details Release / macOS-cpu (arm64, arm64-kleidiai, -DGGML_METAL_USE_BF16=ON -DGGML_METAL_EMBED_LIBRARY=ON -DGGML_CPU_KLEIDIAI=ON, macos-14) (push) Has been cancelled Details Release / macOS-cpu (x64, x64, -DGGML_METAL=OFF -DCMAKE_OSX_DEPLOYMENT_TARGET=13.3, macos-15-intel) (push) Has been cancelled Details Release / ubuntu-24-sycl (fp16, ON) (push) Has been cancelled Details Release / ubuntu-24-sycl (fp32, OFF) (push) Has been cancelled Details Release / windows-hip (gfx1150;gfx1151;gfx1200;gfx1201;gfx1100;gfx1101;gfx1102;gfx1030;gfx1031;gfx1032, radeon) (push) Has been cancelled Details Release / macOS-cpu (arm64, arm64, -DGGML_METAL_USE_BF16=ON -DGGML_METAL_EMBED_LIBRARY=ON, macos-14) (push) Has been cancelled Details Release / ubuntu-cpu (arm64, ubuntu-24.04-arm) (push) Has been cancelled Details Release / ubuntu-cpu (s390x, ubuntu-24.04-s390x) (push) Has been cancelled Details Release / ubuntu-cpu (x64, ubuntu-22.04) (push) Has been cancelled Details Release / ubuntu-vulkan (arm64, ubuntu-24.04-arm) (push) Has been cancelled Details Release / ubuntu-vulkan (x64, ubuntu-22.04) (push) Has been cancelled Details Release / ubuntu-24-openvino (push) Has been cancelled Details Release / windows-cpu (arm64) (push) Has been cancelled Details Release / windows-cpu (x64) (push) Has been cancelled Details Release / windows (arm64, opencl-adreno, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/arm64-windows-llvm.cmake -DCMAKE_PREFIX_PATH="$env:RUNNER_TEMP/opencl-arm64-release" -DGGML_OPENCL=ON -DGGML_OPENCL_USE_ADRENO_KERNELS=ON, ggml-opencl) (push) Has been cancelled Details Release / windows (x64, vulkan, -DGGML_VULKAN=ON, ggml-vulkan) (push) Has been cancelled Details Release / windows-cuda (12.4) (push) Has been cancelled Details Release / windows-cuda (13.1) (push) Has been cancelled Details Release / windows-sycl (push) Has been cancelled Details Release / ios-xcode-build (push) Has been cancelled Details Release / openEuler-cann (aarch64, Release, 310p, off) (push) Has been cancelled Details Release / openEuler-cann (aarch64, Release, 910b, on) (push) Has been cancelled Details Release / openEuler-cann (x86, Release, 310p, off) (push) Has been cancelled Details Release / openEuler-cann (x86, Release, 910b, on) (push) Has been cancelled Details Release / release (push) Has been cancelled Details Release / ui-publish (push) Has been cancelled Details Server (self-hosted) / server-metal (GPUx1, backend-sampling) (push) Has been cancelled Details Server (self-hosted) / server-metal (GPUx2, backend-sampling) (push) Has been cancelled Details Server (self-hosted) / server-metal (GPUx2) (push) Has been cancelled Details Server (self-hosted) / server-metal (GPUx1) (push) Has been cancelled Details Server (self-hosted) / server-kleidiai (CPUx1, kleidiai) (push) Has been cancelled Details Server / server-windows (push) Has been cancelled Details Two fixes surfaced by running the full test suite against the squash-merged turboquant branch, plus one CMake registration. 1. ggml-cuda/ggml-cuda.cu (GET_ROWS supports_op) Removed TQ3_1S/TQ4_1S from the CUDA/HIP GET_ROWS supports_op switch. TheTom's branch advertised these as supported but never added the matching cases to getrows.cu — a latent bug present on both his branch and master. master's test-backend-ops triggers it; the scheduler will now route get_rows on TQ types to CPU. 2. ggml-cuda/fattn.cu (HIP head-size gate) Master's get_best_fattn_kernel falls through to BEST_FATTN_KERNEL_TILE as default. On HIP, fattn-tile.cu only instantiates head sizes 64, 128, 256, 320, 512 (576/640 exceed local memory limits per #ifndef GGML_USE_HIP). Without this gate, supports_op returns true for unsupported sizes and the dispatch aborts. Now returns BEST_FATTN_KERNEL_NONE on HIP for head sizes the tile kernel cannot compile, letting the scheduler fall back to CPU. 3. tests/CMakeLists.txt (test-turbo-quant registration) TheTom added tests/test-turbo-quant.c (CPU round-trip diagnostic for turbo3/turbo4 quant→dequant→inverse-WHT) but never wired it into the build. Registered as a ctest entry linked against ggml + libm. Test status with these fixes: - CPU (build-cpu): 51/51 ctest pass, including new test-turbo-quant. - HIP (build-hip, gfx1151): 50/50 ctest pass with GGML_CUDA_DISABLE_GRAPHS=1 and test-backend-ops excluded. test-backend-ops itself runs 13674/13677 internal cases; the 3 remaining failures (CLAMP f16 → inf, bf16 FA graph capture) are pre-existing master-side regressions on RDNA3.5+HIP that reproduce on plain master and are unrelated to TurboQuant.	2026-05-19 15:13:55 +02:00
shahondin1624	ddebb5ddf6	turboquant: squash-merge TheTom/llama-cpp-turboquant feature/turboquant-kv-cache Squashes the entire TurboQuant KV-cache feature branch from https://github.com/TheTom/llama-cpp-turboquant (tip `5aeb2fdbe`) onto our master. Includes: TurboQuant KV-cache types (turbo2_0, turbo3_0, turbo4_0, tq3_1s, tq4_1s), GGML_OP_TURBO_WHT op, CUDA + Metal kernels (including TQ-rotated mul_mm path), CPU reference paths, HIP template instances, perplexity tooling, and 18 post-upstream-sync fixes (CVE-2026-21869 server clamp, HIP FA pool retention, n_head_v reshape, sparse-V CUDA gating, etc.). Conflict-resolution notes (review carefully before depending on these paths): - common/arg.cpp, common/speculative.cpp: master's refactored speculative API kept (params.speculative.types / ngram_mod struct, per-sinfo n_low/i_last). - ggml-cuda/fattn.cu: head-size exclusion lists unioned (now exclude both 192 and 640 alongside other sizes). - ggml-cuda/ggml-cuda.cu: both master's ADD/SUB/MUL/DIV F16 widening AND TurboQuant's GGML_OP_TURBO_WHT support cases kept. - ggml-metal-device.h/.cpp: master's new get_pipeline_mul_mv_ext signature (const ggml_tensor * op) kept; TurboQuant's get_pipeline_turbo_wht added. - ggml-metal-ops.cpp: TurboQuant's TQ-rotated mul_mm path preserved; non-TQ else-branch adapted to master's pipeline.nr0/nr1/nsg dispatch API. - ggml-vulkan.cpp: master's spec-constant-driven flash_attn pipeline iteration taken (over TurboQuant's CREATE_FA-per-type macro approach). TURBO3_0 added to the fa_kv_ok lambda for type validation. - ggml-vulkan/flash_attn_base.glsl, vulkan-shaders-gen.cpp: master's new spec-constant FA shader generation kept; TurboQuant's DATA_A_TURBO3_0 macro path NOT carried over. * Vulkan TURBO3_0 flash-attention paths need re-implementation against the new spec-constant API. * Vulkan TURBO3_0 inference will likely fail until that work is redone. Squash base: `7fc1c4ef78` (TheTom's last upstream merge point).	2026-05-19 15:13:49 +02:00
Aldehir Rojas	39cf5d6191	common : delegate assistant continuation to underlying template handlers (#23089 ) * common : delegate assistant continuation to template handler * server : implement echo parameter to exclude assistant prefill in the response * server : fix tests for prefill * server : use existing llama template * cont : clean up	2026-05-17 13:36:05 +02:00
Jeff Bolz	7ba22c6a09	vulkan: Support unaligned tensors for ROPE (#22637 )	2026-05-17 11:30:16 +02:00
Aman Gupta	255582687b	llama + spec: MTP Support (#22673 ) * spec: support MTP * fix batch size * rename files * cont : simplify (#7) * MTP: clean-up (#9) * MTP: clean-up * review: use llama_context_type instead of llama_graph_type * review: remove llama_model_has_mtp * review: fix convert issues * convert: fix pycheck * review: formatting * use `mtp-` for identifying mtp models * convert: fix mtp conversion * mtp -> draft-mtp * remove unused llama_arch * add need_embd in speculative * llama: allow partial seq_rm for GDN models for speculative decoding Currently speculative checkpoint needs to restart from a checkpoint after some draft tokens are not accepted, this leads to some wastage in running the target again. This PR adds the ability to rollback upto `draft_max` by storing the GDN intermediates. * fix pending state * vulkan: add GDN partial rollback * meta: extend check to axis 1 * metal: add GDN partial rollback Extend the gated delta net kernel to store intermediate states for partial rollback support on the Metal backend. - Add K (snapshot slot count) as a function constant - Read input state from slot 0 of the 3D state tensor - Write intermediate states to different slots during token loop - For K=1, maintain backward-compatible single-slot behavior Ref: https://github.com/ggml-org/llama.cpp/commit/8c05923630110223669f069af2000e9cf10c02bc Assisted-by: llama.cpp:local pi * delta_net_base: use ggml_pad instead of new_tensor * review: add need_rs_seq * review: rename part_bounded to n_rs * review: deslop comments * review: rename, add asserts * server : adjust checkpoint logic (#11) * server : adjust checkpoint logic * cont : rm asserts * server-context: fix early exit * spec : fix compatibility with n-gram and add TODOs (#13) * metal : cleanup * llama : fix faulty bitwise check in recurrent memory * server : disable RS-based MTP in combination with other spec types * spec : add TODOs * cont : fix comment * cont : update comment * common : fix logic for ngram + mtp compat * llama-memory: enable checkpointing with partial rollback * cont: add test-case for loading into a dirty ctx * llama-memory-recurrent: clear rs_idx in clear * download: fix mtp path * llama-arch: fix enorm op * docs: update docs * conversion: fix type annotations --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2026-05-16 20:06:23 +08:00
Pascal	cfabeb1bad	tests: add BF16 non-contig coverage for MUL_MAT permutations (#22689 ) The MUL_MAT test loop iterates over base_types[] to generate non-contig permutation cases (3 standard permutations across n in {1, 8, 16}). BF16 is absent from base_types[], so these 9 cases were never generated for BF16 even though every other type covered by base_types[] tests them. Add the missing 9 cases explicitly: BF16 x F32, m=16, k=256, bs=[2,3], permutations {0,2,1,3}, {0,1,3,2}, {0,3,2,1}, with n in {1, 8, 16}. Suggested-by: @jeffbolznv	2026-05-15 19:35:05 +02:00
Aman Gupta	ac33f032ac	reasoning-budget: clone should do a deep-copy (#23095 )	2026-05-15 11:59:07 +02:00
Sid Shaytay	91e84fed64	Support for Codex CLI by skipping unsupported Responses tools (#23041 ) * Support for Codex CLI by skipping unsupported Responses tools * Warn on skipped Responses tools and preserve gpt-oss apply_patch rejection * Revert gpt-oss apply_patch special handling	2026-05-15 09:03:24 +02:00
Reese Levine	834a243664	ggml-webgpu: Enable NVIDIA self-hosted CI (#22976 ) * Enabel nvidia ci for webgpu * Address precision issues * fix placement * Relax more set_rows and div * Try relaxing all f16 * formatting and naming * Add comment explaining max_nmse_err logic Added comment referencing pull request for clarification.	2026-05-14 09:41:32 -07:00
Kabir Potdar	42532afff4	unicode,test: add Qwen3.5 non-backtracking tokenizer handler and regr… (#22110 ) * unicode,test: add Qwen3.5 non-backtracking tokenizer handler and regression tests - Add unicode_regex_split_custom_qwen35() to [src/unicode.cpp](src/unicode.cpp), a non-backtracking handler for Qwen3.5's [\p{L}\p{M}]+ regex (letters + combining marks). - Register the handler in the custom tokenizer dispatch table to prevent stack overflows on long inputs (fixes #21919). - Add [models/ggml-vocab-qwen35.gguf](models/ggml-vocab-qwen35.gguf) (test vocab), [models/ggml-vocab-qwen35.gguf.inp](models/ggml-vocab-qwen35.gguf.inp) (test cases), and [models/ggml-vocab-qwen35.gguf.out](models/ggml-vocab-qwen35.gguf.out) (expected output) for regression testing. - Update [tests/CMakeLists.txt](tests/CMakeLists.txt) to include the new test entry. This mirrors the Qwen2 fix (commit `0d049d6`), but adapts for Qwen3.5's regex. Ensures robust Unicode tokenization and prevents std::regex stack overflows. Closes #21919. * fix: enhance regex handling for Qwen3.5 tokenizer to include accent marks * cont : remove trailing whitespace --------- Co-authored-by: Kabir <kabir@example.com> Co-authored-by: Alde Rojas <hello@alde.dev>	2026-05-14 11:03:40 +02:00
Pascal	e936660760	Ggml/cuda snake fusion hardening (#22912 ) * cuda: tighten snake fusion type checks for all operands (defensive, sync vulkan) * cuda: reject snake fusion when ne[2] or ne[3] > 1 (mirror vulkan PR review) * cuda: merge type_ok and types_ok into a single types_ok (address am17an review) * cuda: filter ADD/SUB/MUL/DIV in supports_op to F32/F16 bin_bcast only dispatches F32/F16 type triplets, mirror the vulkan filter so unsupported types fall back through cpy instead of aborting. * test-backend-ops: extend snake_fuse to rank-4 with ne[2]/ne[3] > 1 cases	2026-05-11 18:42:08 +02:00
AesSedai	046e284437	Add flash attention MMA / Tiles to support MiMo-V2.5 (#22812 ) * mimo-v2.5: add flash attention mma/tiles for for d_kq=192 d_v=128 * mimo-v2.5: follow (256, 256) fattn templates * mimo-v2.5: cleanup comments * mimo-v2.5: further comment cleanup * mimo-v2.5: address PR feedback fix GQA handling check for other dangling 320/576 carveouts and mirror them for 192 Add to backend ops test so new paths are covered	2026-05-09 11:28:29 +08:00
Aldehir Rojas	f9cd456ea5	common : revert reasoning budget +inf logit bias (#22740 )	2026-05-08 17:46:43 +02:00
Pascal	58e68df0f9	cuda: fuse snake activation (mul, sin, sqr, mul, add) (#22667 ) * cuda: fuse snake activation (mul, sin, sqr, mul, add) Add ggml_cuda_op_snake_fused with F32 / F16 / BF16 templates. The matcher recognizes the naive 5 op decomposition emitted by audio decoders (BigVGAN, Vocos) for snake activation y = x + sin(ax)^2 inv_b and rewrites it to a single elementwise kernel. Add test_snake_fuse comparing CPU naive vs CUDA fused across F32 / F16 / BF16. * cuda: address review feedback from @am17an Use ggml_cuda_cast for F32/F16/BF16 conversions and rename kernel_snake to snake_kernel to match upstream conventions. * cuda: snake fusion fastdiv on T_len, Suggested-by: @am17an * Update tests/test-backend-ops.cpp Co-authored-by: Aman Gupta <amangupta052@gmail.com> * cuda: snake fusion check add->type matches x->type Address review feedback from @am17an * cuda: snake fusion check add->type matches x->type Moved for readability (equivalent) Address review feedback from @am17an --------- Co-authored-by: Aman Gupta <amangupta052@gmail.com>	2026-05-08 17:44:09 +08:00
leonardHONG	05ff59cb57	CUDA: batch out_prod inner loop with cublasSgemmStridedBatched (#22651 ) * CUDA: batch out_prod inner loop with cublasSgemmStridedBatched * CUDA: batch out_prod inner loop with cublasSgemmStridedBatched * CUDA: add cublasSgemmStridedBatched mapping for HIP and MUSA backends	2026-05-07 21:59:29 +02:00
HaoJun ZHANG	deab41ec68	tests: add long-sequence cases and fix inputs for gated_delta_net (#22794 ) * tests : add long-seq + tail cases for gated_delta_net * tests : realistic input ranges for gated_delta_net	2026-05-08 00:23:36 +08:00
Adrien Gallouët	bf76ac77be	common : only load backends when required (#22290 ) * common : only load backends when required Signed-off-by: Adrien Gallouët <angt@huggingface.co> * llama : call ggml_backend_load_all() directly from llama_backend_init() Signed-off-by: Adrien Gallouët <angt@huggingface.co> * Add ggml_backend_load_all() where llama_backend_init() is not used Signed-off-by: Adrien Gallouët <angt@huggingface.co> --------- Signed-off-by: Adrien Gallouët <angt@huggingface.co>	2026-05-05 09:23:50 +02:00
Ismail	a817a22bc6	ggml : implement fast walsh-hadamard transform for kv rotation (#21352 ) (#22631 )	2026-05-05 10:05:05 +08:00
Piotr Wilkin (ilintar)	a4701c98f7	common/autoparser: fixes for newline handling / forced tool calls (#22654 ) * chat/autoparser: the fixes * Move optspace() to chat-peg-parser, comment out server tests invalidated due to content now allowed with forced tool calls. * Trim whitespace on apply instead	2026-05-04 13:18:11 +02:00
Jeff Bolz	05e141a6b3	vulkan: Support asymmetric FA in coopmat2 path (#21753 ) * vulkan: Support asymmetric FA in coopmat2 path There has been some recent interest/experimentation with mixed quantization types for FA. I had originally designed the cm2 FA shader with this in mind (because I didn't realize it wasn't supported at the time!), this change adds the missing pieces and enables it. Also support Q1_0 since people have been trying that out (seems crazy, but who knows). We should be able to do similar things in the coopmat1/scalar path, but there's another change open against the scalar path and I don't want to conflict. * reorder cases	2026-05-01 15:28:32 +02:00
Anav Prasad	098705a29e	CUDA: fuse SSM_CONV + ADD(bias) + SILU (#22478 )	2026-04-30 02:39:56 +08:00
Masato Nakasaka	7b95ea5d11	common: Intentionally leak logger instance to fix hanging on Windows (#22273 ) * Changed to leak logger singleton to prevent hanging on Windows * Fix comment * Stopped using static vector Using std::vector will cause g_col to be released before the logger thread exits, causing the logger thread to touch freed memory causing a crash * Change so all logs are output before exit * Added debug logging * added more logging * Added logging * Explicitly free logger to avoid hanging on Win * Reverted to leak logger instance again * Removed debug log and fixed comment * Fixed comment --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2026-04-29 10:58:43 +03:00
Michael Wand	fc2b0053ff	ggml-cuda: Repost of 21896: Blackwell native NVFP4 support (#22196 )	2026-04-29 06:47:42 +08:00
Jillis ter Hove	52e5f0a5c1	common : re-arm reasoning budget after DONE on new <think> (#22323 ) DONE state absorbs all tokens including a new start tag, causing any think blocks after the first to run unbudgeted. Observed on unsloth/Qwen3.6-27B-GGUF which interleaves multiple <think> blocks per response. Fixed by advancing start_matcher in DONE branch and re-arming to COUNTING with a fresh budget on match. Adds regression test (test-reasoning-budget: test 6).	2026-04-28 19:15:36 +02:00
Reese Levine	98bb57916a	ggml-webgpu: fix buffer aliasing for ssm_scan and refactor aliasing logic (#22456 ) * Refactor buffer aliasing to be part of shader lib decisions * cleanup * formatting	2026-04-28 07:27:17 -07:00
Georgi Gerganov	14e733e36f	spec : refactor params (#22397 ) * spec : refactor params * cont : fix * cont : rename "sparam" to "sampling" * cont : add spec params category * cont : add info about removed arguments * cont : skip param length check for spec params * cont : adapt server tests	2026-04-28 09:07:33 +03:00
Igor Rudenko	4414c04b9a	Additional test for common/gemma4 : handle parsing edge cases (#22420 ) * Additional test for common/gemma4 : handle parsing edge cases * Move tests to Gemma 4 test group	2026-04-27 16:36:59 +02:00
Piotr Wilkin (ilintar)	dcad77cc3b	chat: fix handling of space in reasoning markers (#22353 ) * chat: fix handling of space in reasoning markers * fix tests * whitespace	2026-04-25 21:24:13 +02:00
Anav Prasad	86db42e97f	CUDA: fuse relu + sqr (#22249 )	2026-04-23 10:28:56 +08:00
Piotr Wilkin (ilintar)	134d6e54d4	common/chat, server: refactor, move all conversion functions to common, add tests (#20690 ) * Refactor conversion functions	2026-04-22 10:28:45 +02:00
Kwa Jie Hao	98d2d2884e	mtmd: Add support for Reka Edge 2603 (#21616 ) * feat: (vocab) fix stray text appended in llama_decode_text Remove accidental concatenation of the full `text` string when formatting UNK_BYTE hex escapes. Only the closing "]" should be appended. * feat(mtmd): add Yasa2 vision encoder support Add a Yasa2 (ConvNeXtV2-based) vision encoder for reka-edge: - Register PROJECTOR_TYPE_YASA2 and tensor name definitions - Add yasa2_block/yasa2_stage model structs - Implement graph builder with ConvNeXt stages, GRN, adaptive pooling - Wire into clip.cpp switch statements and mtmd.cpp init_vision - Use mtmd_image_preprocessor_fixed_size for image preprocessing * feat(chat): add reka-edge template handler (tools, thinking) - Add chat-reka.cpp/h implementing PEG-based parser for reka-edge format - Add Reka-Edge.jinja chat template - Detect reka-edge template in try_specialized_template() - Add LLAMA_EXAMPLE_MTMD to chat-template-file arg * feat: add reka vlm to gguf conversion script Converts Reka Yasa2 hf checkpoints to GGUF format: - Text decoder: Llama-arch with tiktoken/BPE vocab - Mmproj (--mmproj): ConvNeXt vision backbone + language_projection - Generates 2D sincos positional embeddings for vision encoder * test: add Reka Edge chat template and parser tests - test-chat-template: oracle tests comparing Jinja engine output vs common_chat_templates_apply for text, tools, thinking, images, video - test-chat: PEG parser tests for Reka Edge format, round-trip tests for image/video content parts, common path integration tests * scripts: add Reka Edge mixed quantization helper Q4_0 base quantization with Q8_0 override for the last 8 transformer blocks (layers 24-31) via --tensor-type regex. * fix: adapt chat-reka and tests to upstream API - Use autoparser::generation_params (not templates_params) - Add p.prefix(generation_prompt) to PEG parser - Simplify reasoning parser to match LFM2 pattern - Remove image/video oracle tests (unsupported by oaicompat parser; no other multimodal models test this path) * fix: avoid duplicate tensor loading in yasa2 vision encoder TN_YASA_PATCH_W and TN_PATCH_EMBD both resolve to "v.patch_embd.weight", causing the same tensor to be loaded twice into ctx_data and overflowing the memory pool. Reuse the tensors already loaded by the common section. * chore: update image pre-processing settings The reka-edge model depends on the following settings in an older fork of llama.cpp: 1. Fixed square resize 2. BICUBIC 3. add_padding=false In current llama.cpp, this means setting: - image_resize_algo = RESIZE_ALGO_BICUBIC - image_resize_pad = false * chore: remove reka gguf conversion script * chore: remove reka quantization script * chore: remove unnecessary changes from PR scope This commit removes a couple of unnecessary changes for the PR scope: 1. BPE decoder bug fix - this affects reka edge because there's a bug in our tokenization that doesn't represent <think> tokens as special tokens. However this isn't meant to be a thinking model so when run with --reasoning off the edge case does not affect us 2. --chat-template-file support from llama-mtmd-cli - the focus is on llama-server and the reka edge gguf contains the necessary metadata to detect the chat template 3. reka edge oracle test cases - no other model has similar test cases, so I removed it for standardization * chore: remove unnecessary ggml_cast This commit removes unnecessary ggml_cast after updating the reka vlm -> gguf conversion script on hugging face. * chore: remove redundant code * chore: remove unnecessary ggml_cont calls This commit removes all ggml_cont calls except the four that precede ggml_reshape_3d/ggml_reshape_4d. Those are necessary because ggml_reshape recomputes strides assuming contiguous layout and asserts ggml_is_contiguous. Other operations (ggml_mean, ggml_add, ggml_mul etc.) use stride-based indexing and handle non-contiguous inputs correctly and so we are ok to remove ggml_cont for those. * chore: remove unnecessary ggml_repeat calls This commit removes unnecessary ggml_repeat calls because the underlying ops already broadcast automatically. Every ggml_repeat in yasa2.cpp was expanding a smaller tensor to match a larger one's shape before passing both to an elementwise op (ggml_add, ggml_sub, ggml_mul, or ggml_div). This is unnecessary because all four of these ops already support broadcasting internally. * chore: restore ggml_cont needed for cpu operations * refactor: locate reka chat template handler in chat.cpp * chore: remove unnecessary warmup tokens * chore: add code comments on image_resize_pad * chore: remove custom reka parsing code * chore: revert common/chat.cpp * Uncomment debug logging for PEG input parsing --------- Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com>	2026-04-21 20:02:49 +02:00
Aldehir Rojas	d5b780a676	common/autoparser : allow space after tool call (#22073 )	2026-04-19 13:28:35 +02:00
Xuan-Son Nguyen	19124078be	mtmd: add pos_0 to mtmd_image_tokens_get_decoder_pos (breaking change) (#22082 ) * mtmd: add pos_0 to mtmd_image_tokens_get_decoder_pos * fix build	2026-04-19 11:57:21 +02:00
Georgi Gerganov	6990e2f1f7	libs : rename libcommon -> libllama-common (#21936 ) * cmake : allow libcommon to be shared * cmake : rename libcommon to libllama-common * cont : set -fPIC for httplib * cont : export all symbols * cont : fix build_info exports * libs : add libllama-common-base * log : add common_log_get_verbosity_thold()	2026-04-17 11:11:46 +03:00
Piotr Wilkin (ilintar)	e1a9a6dcbe	autoparser: support case of JSON_NATIVE with per-call markers (test case: Reka-Edge) (#21892 )	2026-04-15 10:51:50 +02:00
Xuan-Son Nguyen	707c0b7a6e	mtmd: add mtmd_image_tokens_get_decoder_pos() API (#21851 ) * mtmd: add mtmd_image_tokens_get_decoder_pos() API * consistent naming * fix build	2026-04-14 16:07:41 +02:00
Seyoung Jeong	aa0f1897b7	metal : add XIELU unary op (#20802 )	2026-04-14 15:43:59 +03:00
Aldehir Rojas	e21cdc11a0	common/gemma4 : handle parsing edge cases (#21760 )	2026-04-13 18:18:18 -05:00
Piotr Wilkin (ilintar)	1c0d9081fd	chat: dedicated DeepSeek v3.2 parser + "official" template (#21785 )	2026-04-13 22:23:53 +02:00
Ruben Ortlam	75f3bc94e6	vulkan: Flash Attention DP4A shader for quantized KV cache (#20797 ) * use integer dot product for quantized KV flash attention * small improvements * fix SHMEM_STAGING indexing * add missing KV type quants * fixes * add supported quants to FA tests * readd fast paths for <8bit quants * fix mmq gate and shmem checks	2026-04-13 14:21:31 +02:00
Oliver Simons	9f5e1edb10	CUDA: Limit DeviceSegmentedSort to immediate mode (#21718 ) * CUDA: Limit DeviceSegmentedSort to immediate mode DeviceSegmentedSort is currently not capturable in a cuda graph. Hence, we have to go for the slower DeviceSegmentedRadixSort in that case. Perf numbers on RTX Pro 6000 Blackwell Max-Q: DeviceSegmentedRadixSort in graph mode (i.e. CUDA Graphs) ARGSORT(type=f32,ne=[2048,512,1,1],order=1): 12291 runs - 105.94 us/run - 8192 kB/run - 73.75 GB/s ARGSORT(type=f32,ne=[4096,512,1,1],order=1): 10245 runs - 115.08 us/run - 16384 kB/run - 135.77 GB/s ARGSORT(type=f32,ne=[8192,512,1,1],order=1): 5125 runs - 221.22 us/run - 32768 kB/run - 141.26 GB/s ARGSORT(type=f32,ne=[16384,512,1,1],order=1): 2565 runs - 430.98 us/run - 65536 kB/run - 145.02 GB/s ARGSORT(type=f32,ne=[32768,512,1,1],order=1): 1028 runs - 1185.83 us/run - 131072 kB/run - 105.41 GB/s ARGSORT(type=f32,ne=[65536,512,1,1],order=1): 387 runs - 2748.62 us/run - 262144 kB/run - 90.95 GB/s DeviceSegmentedSort in immediate mode ARGSORT(type=f32,ne=[2048,512,1,1],order=1): 16388 runs - 71.17 us/run - 8192 kB/run - 109.78 GB/s ARGSORT(type=f32,ne=[4096,512,1,1],order=1): 12294 runs - 81.38 us/run - 16384 kB/run - 192.00 GB/s ARGSORT(type=f32,ne=[8192,512,1,1],order=1): 5125 runs - 240.81 us/run - 32768 kB/run - 129.77 GB/s ARGSORT(type=f32,ne=[16384,512,1,1],order=1): 2565 runs - 406.60 us/run - 65536 kB/run - 153.71 GB/s ARGSORT(type=f32,ne=[32768,512,1,1],order=1): 1285 runs - 873.23 us/run - 131072 kB/run - 143.15 GB/s ARGSORT(type=f32,ne=[65536,512,1,1],order=1): 516 runs - 2288.46 us/run - 262144 kB/run - 109.24 GB/s * Add test case for dispatch to DeviceSegmentedRadixSort We currently lack a way to force graph mode in CUDA, patch callback to invoke ggml_backend_compare_graph_backend twice to enforce each test to run in graph mode	2026-04-13 11:14:06 +02:00
Stephen Cox	547765a93e	mtmd: add Gemma 4 audio conformer encoder support (#21421 ) * mtmd: add Gemma 4 audio conformer encoder support Add audio processing for Gemma 4 E2B/E4B via a USM-style Conformer. Architecture: - 12-layer Conformer: FFN → Self-Attention → Causal Conv1D → FFN → Norm - Subsampling Conv Projection: 2x Conv2D(stride=2) with LayerNorm - Full self-attention with sinusoidal RPE and sliding window mask (24) - Logit softcapping at 50.0, ClippableLinear clamping - Output: 1024 → 1536 → RMSNorm → multimodal embedder Mel preprocessing (dedicated mtmd_audio_preprocessor_gemma4a): - HTK mel scale, 128 bins, magnitude STFT, mel_floor=1e-3 - Standard periodic Hann window (320 samples), zero-padded to FFT size - Semicausal left-padding (frame_length/2 samples) - Frame count matched to PyTorch (unfold formula) - No pre-emphasis, no Whisper-style normalization - Mel cosine similarity vs PyTorch: 0.9998 Key fixes: - Tensor loading dedup: prevent get_tensor() from creating duplicate entries in ctx_data. Fixed with std::set guard. - ClippableLinear clamp_info loading moved after per-layer tensors. - Sliding window mask (24 positions) matching PyTorch context_size. - Skip Whisper normalization for Gemma4 mel output. Tested on E2B and E4B with CPU and Vulkan backends. Transcribes: "Glad to see things are going well and business is starting to pick up" (matching ground truth). Ref: #21325	2026-04-12 14:15:26 +02:00
Berk Idem	d7ff074c87	common : enable reasoning budget sampler for gemma4 (#21697 ) * fix: enable reasoning budget sampler for gemma4 Add thinking_start_tag and thinking_end_tag to common_chat_params_init_gemma4(). Without these, the reasoning budget sampler never activates for gemma4. Make the newline after "thought" optional in the PEG parser to handle budget=0 (sampler forces end tag before the newline). Add test case for empty thinking block. Fixes #21487 * use p.space() instead of p.optional(p.literal("\n")) in gemma4 thought parser	2026-04-10 11:49:14 +02:00
Jeff Bolz	7b69125331	vulkan: Support Q1_0 (#21539 ) * vulkan: Support Q1_0 * use get_dm	2026-04-10 08:35:27 +02:00
Johannes Gäßler	d6f3030047	ggml: backend-agnostic tensor parallelism (experimental) (#19378 ) * ggml: backend-agnostic tensor parallelism * support for GPT-OSS, Qwen 3 MoE * partial Vulkan fix * add support for 4/8 GPUs * unconditional peer access * re-use buffers + ggml contexts * fix output pattern * NCCL support * GGML: HIP: add RCCL support * Remove shfl and AllReduce from backend interface * move allocation workaround out of ggml-alloc.c * 2d tensor set/get support * Fix the seg fault without NCCL * Apply suggestion from JohannesGaessler * support for tensor dims % n_devs != 0 * fix view_offs scaling * arbitrary num. of GPUs/tensor split * fix compilation * better granularity estimate * Support device-specific host buffer types if all underlying backends expose the same type. This allows using pinned memory instead of pageable memory for CUDA. Fix compilation errors. * partial Qwen 3 Next support * Fix qwen3 30b (#8) * Fix crash with Qwen-30B-A3B Q4_0 Qwen-30B-A3B Q4_0 has an intermediate dimension of 768. Using a granularity of 256 forces an uneven split between GPUs, which is not supported by the current implementation. * Decide block size based on tensor quantization type * Fix crashes due to KV cache serialization (#9) KV cache serialization requires non-zero offsets on the tensor. Add support in the meta backend to set/get a tensor with a non-zero offset. * metal : fix build (#7) * static memory allocations, fix usage count * fix tensor granularity * more even memory distribution * use BF16 for allreduce * rebase fixup * better error message for unsupported architectures * Fix device mismatch during scatter of allReduce. (#11) There is a mismatch between the dst buffer device and the backend device, causing the use of sync copies * Enable the previous allreduce implementation. It is better in both perf and stability (#12) * delay AllReduce for Moe for less I/O * build : clean-up compile warnings * backend : move most of the meta backend API to ggml-backend-impl.h * cont : hide unused public API in the implementation * llama : use llama_device + remove ggml_backend_dev_is_meta() * ggml-backend : remove unused alloc include * minor : remove regex include * ggml : introduce ggml-ext.h for staging new APIs * rebase fixup * fix tests * llama : more robust logic for determining Meta devices (#16) * llama : more robust logic for determining Meta devices * cont : fix devs size check Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * cont : fix log type Co-authored-by: Johannes Gäßler <johannesg@5d6.de> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * disable roundtrip for meta backend * fix arch selection * Qwen 3.5 support * fix Gemma 4 MoE * fix OpenVino, SYCL * fix test-llama-archs for CPU-only builds * Fix Qwen 3.5 MoE * disable meta backend tests for WebGPU * tests : filter CPU-based devices from the Meta backend tests (#17) * meta : formatting, naming, indentation (#18) * formatting : llama-model.cpp * formatting : ggml-ext.h * formatting : ggml-backend-meta.cpp * meta : add TODO * add documentation * better error messages * fix GPT-OSS --------- Co-authored-by: Carl Philipp Klemm <carl@uvos.xyz> Co-authored-by: Gaurav Garg <gaugarg@nvidia.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2026-04-09 16:42:19 +02:00
Daniel Bevenius	c8ac02fa1b	requirements : update transformers to 5.5.1 (#21617 ) * requirements : update transformers to 5.5.0 This commit updates the transformers dependency to version 5.5.0. The motivation for this is that transformers 5.5.0 includes support for Gemma4 and is required to be able to convert Gemma4 models. This is also causing issues for user of gguf-my-repo. Refs: https://huggingface.co/spaces/ggml-org/gguf-my-repo/discussions/202 * fix huggingface_hub version * set version of transformers to 5.5.0 * convert : add ty ignore directives to convert_hf_to_gguf.py This commit adds `ty: ignore` directives to transformers tokenizers field/methods to avoid type check errors. There might be better ways to handle this and perhaps this can be done in a follow up commit. The motivation for this is that it looks like in transformers 5.5.0 AutoTokenizer.from_pretrained can return generic tokenizer types or None and the type checker now produces an error when the conversion script accesses field like tokenizer.vocab. * convert : add ty ignore to suppress type check errors * convert : remove incorrect type ignores * convert : fix remaining python checks I was running a newer version of ty locally but I've switched to version 0.0.26 which is what CI uses and I was then able to reproduce the errors. Sorry about the noise. * update transformers version to 5.5.1	2026-04-09 12:36:29 +02:00
Piotr Wilkin (ilintar)	0ec191e1d7	vocab: add gemma4 tokenizer tests, fix edge case (#21534 ) * YATF (Yet Another Tokenizer Fix) for Gemma 4. With tests! * Remove unnecessary hash from update script. * minor: move constant	2026-04-09 11:41:14 +02:00
Kwa Jie Hao	243532e556	jinja : support ensure_ascii=true, string repetition and int/float self-filtering (#21623 ) * feat: jinja engine improvements for reka-edge Port three Jinja engine improvements needed for the reka-edge model: 1. Python-style string repetition ("ab" * 3 → "ababab") 2. ensure_ascii=true support for tojson filter (escapes non-ASCII to \uXXXX) 3. int() builtin on value_int_t (identity, needed for Reka Edge template) * fix: escape invalid utf8 bytes when ensure_ascii=true The json_ensure_ascii_preserving_format function does not correctly handle an edge case where if UTF-8 parsing fails, it adds the non-ascii character back to the output as a raw byte. This commit fixes that by adding the unicode standard replacement character \\ufffd to the output instead. This is the standard behavior for various programming languages like Python, Rust, Go, etc. * chore: address PR comments 1. Add todo comment for supporting string repetition for array/tuples 2. Add support for float identity operation 3. Move invalid ascii test case to test_fuzzing * chore: accept suggestion for common/jinja/value.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2026-04-09 11:28:33 +02:00
John Eismeier	e9fd96283d	Propose fix a couple of typos (#21581 ) Signed-off-by: John E <jeis4wpi@outlook.com>	2026-04-08 16:29:03 +02:00
Pasha Khosravi	dcdcbad42a	metal: Q1_0 backend (#21528 ) * initial Q1_0 Metal backend * tuning q1_0 metal kernels * add Q1_0 to test-backend-ops * add Q1_0<->F32 copy test * Apply suggestions from code review Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2026-04-08 16:07:47 +03:00

1 2 3 4 5 ...

777 Commits