发布

  • fix(distributed): track in-flight for non-LLM inference methods (VAD, diarize, voice, ...) (#10238)

    frostbyte_neo 发布于 2026-06-10 14:29:50 +00:00

    fix(distributed): track in-flight for non-LLM inference methods

    InFlightTrackingClient only wrapped a subset of the grpc.Backend
    inference methods (Predict, Embeddings, TTS, AudioTranscription, Detect,
    Rerank, ...). Methods like VAD were left as embedded passthrough, so
    track() never ran for them.

    In distributed mode every model is loaded with in_flight=1 as a
    reservation; that reservation is only released by the OnFirstComplete
    callback, which fires after the first tracked inference call completes.
    A VAD-only model (e.g. silero-vad) never calls a tracked method, so the
    reservation is never released and in-flight stays pinned at 1 forever -
    which also blocks the router's idle-eviction logic.

    Wrap the remaining unary inference methods (VAD, Diarize, Face*, Voice*,
    TokenClassify, Score, AudioEncode, AudioDecode, AudioTransform) with the
    same track()/reconcile() pattern. The three bidi-stream constructors
    (AudioTransformStream, AudioToAudioStream, Forward) are deliberately left
    as passthrough - their inference spans the stream lifetime, not the
    constructor call, so track() there would fire onFirstComplete before any
    data flows.

    Assisted-by: Claude:claude-opus-4-8 [Claude Code]

    Signed-off-by: Ettore Di Giacinto mudler@localai.io
    Co-authored-by: Ettore Di Giacinto mudler@localai.io

    下载附件