server/text_generation_server/layers/attention/rocm.py · 47447ef017d6cdf205be795c7cf7f1086367aa24 · OpenDAS / text-generation-inference

Unify attention output handling (#2343) · 47447ef0

Daniël de Kok authored Aug 01, 2024

- Always return the hidden states.
- Create the output tensor inside the `attention` and `paged_attention`
  functions.

This removes the difference between how the output is handled between
attention (output parameter) and paged attention (return value). This
also removes the assumption that the attention implementation can
write to an output tensor (in preparation of FlashInfer).

47447ef0

rocm.py 7.09 KB

Replace rocm.py