[docs] include bp link. (#11952)

* include bp link. * Update docs/source/en/optimization/fp16.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * resources. --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

[docs] include bp link. (#11952)
* include bp link. * Update docs/source/en/optimization/fp16.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * resources. --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>
5dc503aa · Sayak Paul · GitHub · c6fbcf71 · 5dc503aa
Unverified Commit 5dc503aa authored Jul 18, 2025 by Sayak Paul Committed by GitHub Jul 18, 2025
Show whitespace changes
Inline Side-by-side

Showing with 9 additions and 1 deletion

docs/source/en/optimization/fp16.md docs/source/en/optimization/fp16.md +9 -1

No files found.
--- a/docs/source/en/optimization/fp16.md
+++ b/docs/source/en/optimization/fp16.md
@@ -239,6 +239,12 @@ The `step()` function is [called](https://github.com/huggingface/diffusers/blob/
 In general, the `sigmas` should [stay on the CPU](https://github.com/huggingface/diffusers/blob/35a969d297cba69110d175ee79c59312b9f49e1e/src/diffusers/schedulers/scheduling_euler_discrete.py#L240) to avoid the communication sync and latency.
+<Tip>
+Refer to the [torch.compile and Diffusers: A Hands-On Guide to Peak Performance](https://pytorch.org/blog/torch-compile-and-diffusers-a-hands-on-guide-to-peak-performance/) blog post for maximizing performance with `torch.compile` for diffusion models.
+</Tip>
 ### Benchmarks
 Refer to the [diffusers/benchmarks](https://huggingface.co/datasets/diffusers/benchmarks) dataset to see inference latency and memory usage data for compiled pipelines.
@@ -299,3 +305,5 @@ pipeline.fuse_qkv_projections()
 - Read the [Presenting Flux Fast: Making Flux go brrr on H100s](https://pytorch.org/blog/presenting-flux-fast-making-flux-go-brrr-on-h100s/) blog post to learn more about how you can combine all of these optimizations with [TorchInductor](https://docs.pytorch.org/docs/stable/torch.compiler.html) and [AOTInductor](https://docs.pytorch.org/docs/stable/torch.compiler_aot_inductor.html) for a ~2.5x speedup using recipes from [flux-fast](https://github.com/huggingface/flux-fast).
    These recipes support AMD hardware and [Flux.1 Kontext Dev](https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev).
+- Read the [torch.compile and Diffusers: A Hands-On Guide to Peak Performance](https://pytorch.org/blog/torch-compile-and-diffusers-a-hands-on-guide-to-peak-performance/) blog post
+to maximize performance when using `torch.compile`.
\ No newline at end of file