-
feat(llama-cpp): bump to `1ec7ba0c`, adapt grpc-server, expose new spec-decoding options (#9765)
发布于
2026-05-12 15:22:37 +00:00 - chore(llama.cpp): bump to 1ec7ba0c14f33f17e980daeeda5f35b225d41994
Picks up the upstream
spec : parallel drafting supportchange
(ggml-org/llama.cpp#22838) which reshapes the speculative-decoding API
andserver_context_impl.Adapt the grpc-server wrapper accordingly:
common_params_speculative::type(single enum) becametypes
(std::vector<common_speculative_type>). Update both the
"default to draft when a draft model is set" branch and the
spec_type/speculative_typeoption parser. The parser now also
tolerates comma-separated lists, mirroring the upstream
common_speculative_types_from_namessemantics.common_params_speculative_draft::n_ctxis gone (draft now shares
the target context size). Keep thedraft_ctx_sizeoption name for
backward compatibility and ignore the value rather than failing.server_context_impl::modelwas renamed tomodel_tgt; update the
two reranker / model-metadata call sites.
Replaces #9763. Builds cleanly under the linux/amd64 cpu-llama-cpp
target locally.Signed-off-by: Ettore Di Giacinto mudler@localai.io
- feat(llama-cpp): expose new speculative-decoding option keys
Upstream
spec : parallel drafting support(ggml-org/llama.cpp#22838)
adds thengram_mod,ngram_map_k, andngram_map_k4vspeculative
families and beefs up the draft-model knobs. The previous bump only
adapted the API; this exposes the new fields through the grpc-server
options dictionary so model configs can drive them.New
options:keys (all underbackend: llama-cpp):ngram_mod (
ngram_modtype):
spec_ngram_mod_n_min / spec_ngram_mod_n_max / spec_ngram_mod_n_matchngram_map_k (
ngram_map_ktype):
spec_ngram_map_k_size_n / spec_ngram_map_k_size_m / spec_ngram_map_k_min_hitsngram_map_k4v (
ngram_map_k4vtype):
spec_ngram_map_k4v_size_n / spec_ngram_map_k4v_size_m /
spec_ngram_map_k4v_min_hitsngram lookup caches (
ngram_cachetype):
spec_lookup_cache_static / lookup_cache_static
spec_lookup_cache_dynamic / lookup_cache_dynamicDraft-model tuning (active when
spec_typeisdraft):
draft_cache_type_k / spec_draft_cache_type_k
draft_cache_type_v / spec_draft_cache_type_v
draft_threads / spec_draft_threads
draft_threads_batch / spec_draft_threads_batch
draft_cpu_moe / spec_draft_cpu_moe (bool flag)
draft_n_cpu_moe / spec_draft_n_cpu_moe (first N MoE layers on CPU)
draft_override_tensor / spec_draft_override_tensor
(comma-separated =; re-implements upstream's
static parse_tensor_buffer_overrides since it isn't exported)spec_typealready accepted comma-separated lists after the previous
commit, matching upstream'scommon_speculative_types_from_names.Docs: refresh
docs/content/advanced/model-configuration.mdwith
per-family tables and a note about multi-type chaining.Builds locally with
make docker-build-llama-cpp(linux/amd64
cpu-llama-cpp AVX variant).Signed-off-by: Ettore Di Giacinto mudler@localai.io
- fix(turboquant): bridge new llama.cpp spec API to the legacy fork layout
The previous commits in this series adapted backend/cpp/llama-cpp/grpc-server.cpp
to the post-#22838 (parallel drafting) llama.cpp API. The turboquant build
reuses the same grpc-server.cpp through backend/cpp/turboquant/Makefile,
which copies it into turboquant--build/ and runs patch-grpc-server.sh
on the copy. The fork branched before the API refactor, so it errors out on:ctx_server.impl->model_tgt(fork still hasmodel)params.speculative.{ngram_mod,ngram_map_k,ngram_map_k4v,ngram_cache}.*
(none of these sub-structs exist in the fork)params.speculative.draft.{cache_type_k/v, cpuparams[, _batch].n_threads, tensor_buft_overrides}(fork uses the pre-#22397 flat layout)params.speculative.typesvector /common_speculative_types_from_names
(fork has a scalartypeand only the singular helper)
Approach:
-
backend/cpp/llama-cpp/grpc-server.cpp: introduce a single feature switch
LOCALAI_LEGACY_LLAMA_CPP_SPEC. When defined, the twospeculative.type[s]
discriminations (the "default to draft when a draft model is set" branch
and thespec_type/speculative_typeoption parser) fall back to the
singular scalar form, and the entire new-option block (ngram_mod / map_k
/ map_k4v / ngram_cache / draft.{cache_type_, cpuparams,
tensor_buft_overrides}) is preprocessed out. The macro is not defined
in the source tree — stock llama-cpp builds get the full new API. -
backend/cpp/turboquant/patch-grpc-server.sh: two new patch steps applied
to the per-flavor build copy at turboquant--build/grpc-server.cpp:- substitute
ctx_server.impl->model_tgt->ctx_server.impl->model - inject
#define LOCALAI_LEGACY_LLAMA_CPP_SPEC 1before the first
#include, so the guarded blocks above drop out for the fork build.
Both patches are idempotent and follow the existing sed/awk pattern in
this script (KV cache types,get_media_marker, flat speculative
renames). Stock llama-cpp'sgrpc-server.cppis never touched. - substitute
Drop both legacy patches once the turboquant fork rebases past
ggml-org/llama.cpp#22397 / #22838.Signed-off-by: Ettore Di Giacinto mudler@localai.io
- fix(turboquant): close draft_ctx_size brace inside legacy guard
The previous turboquant fix wrapped the new option-handler blocks in
#ifndef LOCALAI_LEGACY_LLAMA_CPP_SPEC ... #endifbut placed the guard
in the middle of anelse ifchain — the} else ifopenings of the
new blocks were responsible for closing the previous block's brace.
With the macro defined the new blocks vanish, draft_ctx_size's{
loses its closer, the for-loop's}is consumed instead, and the
file ends with a stray opening brace — clang reports it as
function-definition is not allowed here before '{'on the next
top-levelint main(...)andexpected '}' at end of input.Move the chain split inside the draft_ctx_size branch:
} else if (... "draft_ctx_size") { // ...#ifdef LOCALAI_LEGACY_LLAMA_CPP_SPEC
} // legacy: chain ends here
#else
} else if (... "spec_ngram_mod_n_min") { // modern: chain continues
...
} else if (... "draft_override_tensor") {
...
} // closes last branch
#endif
} // closes for-loopBrace count is now balanced under both preprocessor branches (verified
withtr -cd '{' | wc -cagainst the patched and unpatched outputs).Local
make docker-build-turboquantbuilds the linux/amd64 cpu-llama-cpp
turboquant-avxvariant cleanly.Signed-off-by: Ettore Di Giacinto mudler@localai.io
- fix(ci): forward AMDGPU_TARGETS into Dockerfile.turboquant builder-prebuilt
Dockerfile.turboquant's
builder-prebuiltstage was missing the
ARG AMDGPU_TARGETS/ENV AMDGPU_TARGETS=${AMDGPU_TARGETS}pair that
builder-fromsourcealready has (and thatDockerfile.llama-cpp
mirrors across both stages). When CI uses the prebuilt base image
(quay.io/go-skynet/ci-cache:base-grpc-*, the common path) the build-arg
passed by the workflow never reaches the env inside the compile stage.backend/cpp/llama-cpp/Makefile:38 (introduced by #9626) errors out on
hipblas builds when AMDGPU_TARGETS is empty, and the turboquant
Makefile reuses backend/cpp/llama-cpp via a sibling build dir, so the
same check fires from turboquant-fallback under BUILD_TYPE=hipblas:Makefile:38: *** AMDGPU_TARGETS is empty — set it to a comma-separated
list of gfx targets e.g. gfx1100,gfx1101. Stop.
make: *** [Makefile:66: turboquant-fallback] Error 2The bug is latent on master because the docker layer cache stays warm
across builds — the compile step rarely re-runs from scratch. The
llama.cpp bump in this PR invalidates the cache, so the missing env var
becomes load-bearing and the hipblas turboquant CI job fails.Mirror the existing pattern from Dockerfile.llama-cpp.
Signed-off-by: Ettore Di Giacinto mudler@localai.io
Signed-off-by: Ettore Di Giacinto mudler@localai.io
Co-authored-by: Ettore Di Giacinto mudler@localai.io下载附件