Fix typo in "attention" (#7977)

d6ca1209 · Jacob Marks · GitHub · fb7ae018 · d6ca1209
Unverified Commit d6ca1209 authored May 20, 2024 by Jacob Marks Committed by GitHub May 20, 2024
Show whitespace changes
Inline Side-by-side

Showing with 1 addition and 1 deletion

docs/source/en/optimization/fp16.md docs/source/en/optimization/fp16.md +1 -1

No files found.
--- a/docs/source/en/optimization/fp16.md
+++ b/docs/source/en/optimization/fp16.md
@@ -12,7 +12,7 @@ specific language governing permissions and limitations under the License.

 # Speed up inference

-There are several ways to optimize Diffusers for inference speed, such as reducing the computational burden by lowering the data precision or using a lightweight distilled model. There are also memory-efficient attention implementations, [xFormers](xformers) and [scaled dot product attetntion](https://pytorch.org/docs/stable/generated/torch.nn.functional.scaled_dot_product_attention.html) in PyTorch 2.0, that reduce memory usage which also indirectly speeds up inference. Different speed optimizations can be stacked together to get the fastest inference times.
+There are several ways to optimize Diffusers for inference speed, such as reducing the computational burden by lowering the data precision or using a lightweight distilled model. There are also memory-efficient attention implementations, [xFormers](xformers) and [scaled dot product attention](https://pytorch.org/docs/stable/generated/torch.nn.functional.scaled_dot_product_attention.html) in PyTorch 2.0, that reduce memory usage which also indirectly speeds up inference. Different speed optimizations can be stacked together to get the fastest inference times.

 > [!TIP]
 > Optimizing for inference speed or reduced memory usage can lead to improved performance in the other category, so you should try to optimize for both whenever you can. This guide focuses on inference speed, but you can learn more about lowering memory usage in the [Reduce memory usage](memory) guide.