README.md

# HunyuanDiT

## 论文

**Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding**

* https://arxiv.org/pdf/2405.08748

## 模型结构

模型基于`transformer decoder`结构，在`DiT`基础上重新设计了`Time Embedding`以及`positional Embedding`的添加方式，`Text Prompt`通过两个`text encoder`进行编码，其余与DiT一致。

![alt text](docs/model_structure.png)


## 算法原理
使用`self-attention`捕获图像内部的结构信息，使用`cross attention`对齐文本与图像。

![alt text](docs/alg.png)


## 环境配置

### Docker（方法一）
    
    docker pull image.sourcefind.cn:5000/dcu/admin/base/pytorch:2.1.0-ubuntu20.04-dtk24.04.1-py3.10

    docker run --shm-size 10g --network=host --name=hunyuan --privileged --device=/dev/kfd --device=/dev/dri --group-add video --cap-add=SYS_PTRACE --security-opt seccomp=unconfined -v 项目地址(绝对路径):/home/ -v /opt/hyhal:/opt/hyhal:ro -it <your IMAGE ID> bash

### Dockerfile（方法二）

    docker build -t <IMAGE_NAME>:<TAG> .

    docker run --shm-size 10g --network=host --name=hunyuan --privileged --device=/dev/kfd --device=/dev/dri --group-add video --cap-add=SYS_PTRACE --security-opt seccomp=unconfined -v 项目地址(绝对路径):/home/ -v /opt/hyhal:/opt/hyhal:ro -it <your IMAGE ID> bash

## 数据集

无

## 推理

### 安装diffuser和依赖

```
git clone https://developer.hpccube.com/codes/modelzoo/hunyuandit_diffusers.git
cd hunyuandit_diffusers
git submodule init && git submodule update

1. 卸载apex
pip uninstall apex
2. 安装diffusers
cd diffusers
python3 setup.py install
cd ..

```

### 模型下载

模型快速下载中心：[AIModels](http://113.200.138.88:18080/aimodels), 本项目模型链接：[stable-diffusion-3-medium-diffusers](http://113.200.138.88:18080/aimodels/HunyuanDiT)

### 运行

    export HY_USE_XFORMERS=1
    python SD-HunyuanDit.py

## result

![img](./docs/result.png)

### 精度

无

## 应用场景

### 算法类别

`AIGC`

### 热点应用行业

`零售,广媒,电商`

## 源码仓库及问题反馈

* https://developer.hpccube.com/codes/modelzoo/hunyuandit_diffusers

## 参考资料

* https://hf-mirror.com/Tencent-Hunyuan/HunyuanDiT-Diffusers