1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
1. 视觉 Tokenizer
├── VQ-VAE ← 建议补
└── VQGAN

2. Autoregressive Image Generation
└── DALL·E

3. Diffusion Model
├── 3.1 Diffusion 基础
│ └── DDPM

├── 3.2 Latent Diffusion
│ └── Latent Diffusion Model

└── 3.3 Denoiser Backbone 演进
├── U-Net
└── DiT

4. Flow-based Generation
└── Flow Matching

5. Vision-Language Representation
└── CLIP

6. Multimodal LLM
├── BLIP-2
└── LLaVA