4c4911fe2c
* chore(llama-cpp): update upstream revision Assisted-by: Codex:gpt-5.6 * fix(llama-cpp): refresh server patch contexts The new llama.cpp pin changed the slot reset and prompt batch code. GNU patch accepted stale hunks with fuzz, which left the L4T build with invalid source. Refresh both server patches against the pinned source so each hunk applies at its intended location. Assisted-by: Codex:gpt-5 * fix(llama-cpp): adapt metrics result fields The updated llama.cpp groups cumulative counters under server_metrics. Probe the result layout so the shared adapter also compiles against older forks. Assisted-by: Codex:gpt-5 * fix(llama-cpp): refresh TTS patch offsets GNU patch rejects the stale pre-decode hunk after the score patch changes the same file. Anchor the TTS hunks to the pinned llama.cpp source so the full series applies without fuzz. Assisted-by: Codex:gpt-5.4 * fix(llama-cpp): normalize batch threads The updated llama.cpp creates its batch threadpool during model initialization, before the context-level fallback can replace the -1 sentinel. Resolve that sentinel from the inference thread count so model loading does not overflow the threadpool allocation.\n\nAssisted-by: Codex:gpt-5.4 --------- Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
78 lines
3.2 KiB
Bash
78 lines
3.2 KiB
Bash
#!/bin/bash
|
|
|
|
set -e
|
|
|
|
|
|
## Patches
|
|
|
|
## Apply patches from the `patches` directory. Runs under set -e so a
|
|
## rejected patch aborts the build here, loudly, instead of surfacing later
|
|
## as a confusing compile error. A missing or empty patches dir is a no-op.
|
|
if [ -d "patches" ]; then
|
|
for patch in $(ls patches); do
|
|
echo "Applying patch $patch"
|
|
patch -d llama.cpp/ -p1 < patches/$patch
|
|
done
|
|
fi
|
|
|
|
for file in $(ls llama.cpp/tools/server/); do
|
|
cp -rfv llama.cpp/tools/server/$file llama.cpp/tools/grpc-server/
|
|
done
|
|
|
|
cp -r CMakeLists.txt llama.cpp/tools/grpc-server/
|
|
cp -r grpc-server.cpp llama.cpp/tools/grpc-server/
|
|
# Shared message-reconstruction helpers (included by grpc-server.cpp) and their
|
|
# unit test (compiled only when -DLLAMA_GRPC_BUILD_TESTS=ON).
|
|
cp -r message_content.h llama.cpp/tools/grpc-server/
|
|
cp -r message_content_test.cpp llama.cpp/tools/grpc-server/
|
|
# Generic passthrough parser staging and its standalone regression test.
|
|
cp -r passthrough_options.h llama.cpp/tools/grpc-server/
|
|
cp -r passthrough_options_test.cpp llama.cpp/tools/grpc-server/
|
|
# TTS request validation (included by grpc-server.cpp) and its standalone
|
|
# regression test.
|
|
cp -r tts_request_options.h llama.cpp/tools/grpc-server/
|
|
cp -r tts_request_options_test.cpp llama.cpp/tools/grpc-server/
|
|
# Thread-count default normalization and its standalone regression test.
|
|
cp -r thread_params.h llama.cpp/tools/grpc-server/
|
|
cp -r thread_params_test.cpp llama.cpp/tools/grpc-server/
|
|
# Parent-death watcher (included by grpc-server.cpp) and its standalone unit
|
|
# test (run via backend/cpp/run-unit-tests.sh; also buildable under ctest).
|
|
cp -r parent_watch.h llama.cpp/tools/grpc-server/
|
|
cp -r parent_watch_test.cpp llama.cpp/tools/grpc-server/
|
|
cp -rfv llama.cpp/vendor/nlohmann/json.hpp llama.cpp/tools/grpc-server/
|
|
cp -rfv llama.cpp/vendor/cpp-httplib/httplib.h llama.cpp/tools/grpc-server/
|
|
|
|
## Fork-skew probe. Upstream folded common_params::use_mmap / use_mlock /
|
|
## use_direct_io into a single `load_mode` enum (ggml-org/llama.cpp#20834).
|
|
## turboquant and bonsai compile this very same grpc-server.cpp against forks
|
|
## that branched before that change, so the field set is decided from the
|
|
## checkout in front of us rather than from a per-fork build flag: the flavor
|
|
## targets disagree on whether they forward CMAKE_ARGS or EXTRA_CMAKE_ARGS, and
|
|
## probing heals itself the moment a fork rebases past the refactor.
|
|
if grep -q "LLAMA_LOAD_MODE_MMAP" llama.cpp/include/llama.h; then
|
|
echo "==> llama.cpp carries the load-mode enum, using common_params::load_mode"
|
|
LEGACY_LOAD_MODE=0
|
|
else
|
|
echo "==> llama.cpp predates the load-mode enum, using the legacy mmap/mlock/direct-io booleans"
|
|
LEGACY_LOAD_MODE=1
|
|
fi
|
|
if grep -q "server_metrics metrics;" llama.cpp/tools/server/server-task.h; then
|
|
HAS_SERVER_METRICS=1
|
|
else
|
|
HAS_SERVER_METRICS=0
|
|
fi
|
|
cat > llama.cpp/tools/grpc-server/llama_compat.h <<EOF
|
|
// Generated by backend/cpp/llama-cpp/prepare.sh. Do not edit.
|
|
#pragma once
|
|
#define LOCALAI_LEGACY_LOAD_MODE ${LEGACY_LOAD_MODE}
|
|
#define LOCALAI_HAS_SERVER_METRICS ${HAS_SERVER_METRICS}
|
|
EOF
|
|
|
|
set +e
|
|
if grep -q "grpc-server" llama.cpp/tools/CMakeLists.txt; then
|
|
echo "grpc-server already added"
|
|
else
|
|
echo "add_subdirectory(grpc-server)" >> llama.cpp/tools/CMakeLists.txt
|
|
fi
|
|
set -e
|