1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27
| 1. 视觉 Tokenizer ├── VQ-VAE ← 建议补 └── VQGAN
2. Autoregressive Image Generation └── DALL·E
3. Diffusion Model ├── 3.1 Diffusion 基础 │ └── DDPM │ ├── 3.2 Latent Diffusion │ └── Latent Diffusion Model │ └── 3.3 Denoiser Backbone 演进 ├── U-Net └── DiT
4. Flow-based Generation └── Flow Matching
5. Vision-Language Representation └── CLIP
6. Multimodal LLM ├── BLIP-2 └── LLaVA
|