发布

  • Fix: Adapt Llama injection policy for newer transformers versions (#7443)

    frostbyte_neo 发布于 2025-07-26 21:27:33 +00:00

    This PR fixes an AttributeError that occurs during
    deepspeed.init_inference when using kernel injection
    (replace_with_kernel_inject=True) with Llama models from recent
    versions of transformers.

    The Bug:

    In newer transformers versions (e.g., 4.53.3), configurations like
    num_heads and rope_theta were moved from direct attributes of the
    LlamaAttention module into a nested config object.

    The current DeepSpeed injection policy tries to access these attributes
    from their old, direct location, causing the initialization to fail with
    an AttributeError: 'LlamaAttention' object has no attribute 'num_heads'.

    The Solution:

    This change updates the Llama injection logic to be more robust:

    1. It first tries to read attributes like num_heads from the new
      config object location.
    2. If that fails, it falls back to the legacy direct attribute path.

    Signed-off-by: huanyuqu yc37960@um.edu.mo

    下载附件