目录

1、扩散模型Awesome 汇总

2、扩散模型综述汇总

3、自动驾驶中的扩散模型

4、机器人数据

5、数据集生成


1、扩散模型Awesome 汇总

1、关于扩散模型的资源和论文集

Awesome-Diffusion-Models库是一个与扩散模型相关的资源集合,包括介绍性文章、论文、视频、教程,以及各种应用程序的研究,如图像生成、医学成像和强化学习。它被组织成视觉、音频、自然语言处理等主题。此外,它还包括调查,教程和木星笔记本电脑,以帮助用户开始或加深他们对扩散模型的理解 

https://github.com/diff-usion/Awesome-Diffusion-Models

2、视频生成、编辑、恢复、理解等最新传播模型列表

https://github.com/showlab/Awesome-Video-Diffusion

3、基于扩散的图像处理综述,包括恢复、增强、编码、质量评估

https://github.com/lixinustc/Awesome-diffusion-model-for-image-processing

4、图扩散生成工作集合,包括论文、代码和数据集。

https://github.com/yuntaoshou/Graph-Diffusion-Models-A-Comprehensive-Survey-of-Methods-and-Applications

2、扩散模型综述汇总

1、探索自动驾驶中视频生成与世界模型之间的相互作用:一项调查

    2、扩散模型在3D视觉中的算法及应用全面综述

      3、扩散模型如何在智能交通(自动驾驶、交通仿真、轨迹预测等)领域发挥作用?

        4、扩散模型及其应用全面综述

          5、首个围绕低层次视觉任务中去噪扩散模型技术全面综述

            6、Efficient Diffusion Models: A Comprehensive Survey from Principles to Practices

              7、Diffusion Models in 3D Vision: A Survey

                8、Conditional Image Synthesis with Diffusion Models: A Survey

                  9、Trustworthy Text-to-Image Diffusion Models: A Timely and Focused Survey

                    10、A Survey on Diffusion Models for Recommender Systems

                      11、Diffusion-Based Visual Art Creation: A Survey and New Perspectives

                        12、Replication in Visual Diffusion Models: A Survey and Outlook

                          13、Diffusion Model-Based Video Editing: A Survey

                            14、Diffusion Models and Representation Learning: A Survey

                              15、A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

                                16、Diffusion Models in Low-Level Vision: A Survey

                                  17、Video Diffusion Models: A Survey

                                    18、A Survey on Diffusion Models for Time Series and Spatio-Temporal Data

                                      19、Controllable Generation with Text-to-Image Diffusion Models: A Survey

                                        20、Diffusion Model-Based Image Editing: A Survey

                                          21、Diffusion Models, Image Super-Resolution And Everything: A Survey

                                            22、A Survey on Video Diffusion Models

                                              23、A Survey of Diffusion Models in Natural Language Processing

                                                3、自动驾驶中的扩散模型

                                                1、为自动驾驶应用采集车辆资产!Drive-1-to-3: 丰富扩散先验的实车新视图合成方法

                                                  2、一种新颖的单域目标检测泛化方法——GoDiff

                                                    3、StreetCrafter:一种新型的可控自动驾驶街景合成视频扩散模型

                                                      4、DiffusionDrive:面向端到端自动驾驶的截断扩散模

                                                        5、MagicDriveDiT:基于自适应控制的自动驾驶高分辨率长视频生成

                                                          6、Cityscape-Adverse:利用基于扩散的图像编辑来模拟八种不利条件

                                                            7、DrivingDiffusion:一种新颖的时空一致扩散框架

                                                              8、Diffusion-Occ——一种新颖的点云补全框架

                                                                9、【ECCV 2024】扩散模型都可以用于单目深度估计了?

                                                                  10、OccSora:一种基于扩散的4D占用生成模型,用于模拟自动驾驶3D世界的发展

                                                                    11、3DiffTection:这是一种利用3D感知扩散模型的特征从单个图像中检测3D目标的最先进方法

                                                                      12、OLiDM: Object-aware LiDAR Diffusion Models for Autonomous Driving

                                                                        13、SynDiff-AD: Improving Semantic Segmentation and End-to-End Autonomous Driving with Synthetic Data from Latent Diffusion Models

                                                                          14、Characterized Diffusion Networks for Enhanced Autonomous Driving Trajectory Prediction

                                                                            15、DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

                                                                              16、Efficient Domain Augmentation for Autonomous Driving Testing Using Diffusion Models

                                                                                17、DriveDiTFit: Fine-tuning Diffusion Transformers for Autonomous Driving

                                                                                  18、VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving

                                                                                    19、Enhanced Safety in Autonomous Driving: Integrating Latent State Diffusion Model for End-to-End Navigation

                                                                                      20、Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous Driving and Zero-Shot Instruction Following

                                                                                        21、Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion

                                                                                          4、机器人数据

                                                                                          RoboMIND:上央视新闻啦!我国首个通用多本体具身智能数据集发布

                                                                                          百万机器人真机数据集

                                                                                          GitHub链接:

                                                                                          https://github.com/OpenDriveLab/AgiBot-World

                                                                                          抱抱脸链接:

                                                                                          ‍https://huggingface.co/agibot-world

                                                                                          项目主页:

                                                                                          https://agibot-world.com/

                                                                                          5、数据集生成

                                                                                          1、绘画

                                                                                          https://github.com/poloclub/diffusiondb

                                                                                          2、数据集扩散:基于扩散的像素级语义分割合成数据生成

                                                                                          https://github.com/VinAIResearch/Dataset-Diffusion

                                                                                          3、用于文本到视频扩散模型的百万级实时图库数据集

                                                                                          https://github.com/WangWenhao0716/VidProM

                                                                                          4、全球首个车路协同自动驾驶数据集

                                                                                          https://air.tsinghua.edu.cn/DAIR-V2X/index.html

                                                                                          5、其它CVPR24中的数据集

                                                                                          Title

                                                                                          Authors

                                                                                          Summary

                                                                                          Benchmarking and Evaluating Large Video Generation Models

                                                                                          Yaofang Liu, Xiaodong Cun, Xuebo Liu

                                                                                          The paper proposes a comprehensive evaluation framework for large video generation models that have grown rapidly. Existing academic metrics are inadequate for evaluating these models trained on massive datasets. The proposed evaluation pipeline comprises prompt curation, objective evaluation, subjective studies, and opinion alignment. The models are evaluated based on 17 objective metrics covering visual quality, content quality, motion quality, and text-caption alignment. Additionally, it provides a comparison table of various video generation models across different metrics and capabilities.

                                                                                          ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic Object

                                                                                          Chenshuang Zhang, Fei Pan, Junmo Kim

                                                                                          ImageNet-D is a new benchmark for evaluating neural network robustness in visual perception tasks. It generates synthetic images with diverse backgrounds, textures, and materials, making it more challenging than other synthetic datasets. Key features include diversified image generation, high visual fidelity, and significant accuracy reduction of various vision models. The benchmark is created by combining object categories and refining through human verification. ImageNet-D is effective in evaluating neural network robustness, as accuracy on it improves with accuracy on ImageNet.

                                                                                          Polos: Multimodal Metric Learning from Human Feedback for Image Captioning

                                                                                          Yuiga Wada, Kanta Kaneda, Daichi Saito, Komei Sugiura

                                                                                          The Polaris dataset, used to train the model, contains 131,020 human judgments from 550 evaluators on the appropriateness of image captions. The dataset is much larger than existing ones and is capable of training image captioning metrics. The captions in Polaris are more diverse, collected from humans and generated by 10 modern image captioning models. This demonstrates the effectiveness and robustness of Polos compared to previous metrics.

                                                                                          VBench: Comprehensive Benchmark Suite for Video Generative Models

                                                                                          Ziqi Huang, Yinan He, Jiashuo Yu

                                                                                          VBench is a tool that evaluates video generation models across 16 quality dimensions. These dimensions fall under Video Quality and Video-Condition Consistency. VBench provides valuable insights by evaluating models across multiple dimensions, content categories, and comparing video vs image generation. The tool's authors plan to expand VBench to more models and video generation tasks. Checkout the leaderboard on HF here: https://huggingface.co/spaces/Vchitect/VBench_Leaderboard

                                                                                          Diffusion 3D Features (Diff3F): Decorating Untextured Shapes with Distilled Semantic Features

                                                                                          Niladri Shekhar Dutt, Sanjeev Muralikrishnan, Niloy J. Mitra

                                                                                          Diff3F is a feature descriptor for untextured 3D shapes. It computes 3D semantic features using pre-trained 2D diffusion models, rendering depth and normal maps from multiple views, and lifting the 2D diffusion features back to the 3D surface. This produces semantic descriptors on the 3D shape without requiring additional training data or part segmentation.

                                                                                          One-step Diffusion with Distribution Matching Distillation

                                                                                          Tianwei Yin

                                                                                          Distribution Matching Distillation (DMD accelerates multi-step diffusion models into a one-step generator without compromising image quality. DMD matches the distribution of the original diffusion model by minimizing KL divergence and using two score functions - one for the actual data distribution and one for the generated distribution. A regression loss matches the large-scale structure of the multi-step diffusion outputs.

                                                                                          Logo

                                                                                          北京人形旗下天工造物具身智能开源社区,聚焦具身天工与慧思开物两大平台

                                                                                          更多推荐