# How DALL-E 2 Actually Works

How does OpenAI's groundbreaking DALL-E 2 model actually work? Check out this detailed guide to learn the ins and outs of DALL-E 2.

**DALL-E 3**

OpenAI has recently announced DALL-E 3, the successor to DALL-E 2. For information on what DALL-E 3 is, how it works, and the differences between DALL-E 3 and DALL-E 2, jump down to [**this section**](/content/blog/how-dall-e-2-actually-works/#what-is-dall-e-3/index.html).

OpenAI's groundbreaking model [**DALL-E 2**](https://openai.com/dall-e-2/?ref=assemblyai.com) hit the scene at the beginning of the month, setting a new bar for image generation and manipulation. With only a short text prompt, DALL-E 2 can **generate completely new images** that combine distinct and unrelated objects in semantically plausible ways, like the images below which were generated by entering the prompt **"a bowl of soup that is a portal to another dimension as digital art"**.

DALL-E 2 can even modify existing images, create variations of images that maintain their salient features, and interpolate between two input images. DALL-E 2's impressive results have many wondering exactly how such a powerful model works under the hood.

In this article, **we will take an in-depth look at how DALL-E 2 manages to create such astounding images** like those above. Plenty of background information will be given and the explanation levels will run the gamut, so this article is suitable for readers at several levels of Machine Learning experience. Let's dive in!

## How DALL-E 2 Works: A Bird's-Eye View

Before diving into the details of how DALL-E 2 works, let's orient ourselves with a high-level overview of how DALL-E 2 generates images. While DALL-E 2 can perform a variety of tasks, including image manipulation and interpolation as mentioned above, **we will focus on the task of image generation** in this article.

At the highest level, DALL-E 2's works very simply:

1. First, a text prompt is input into a **text encoder** that is trained to map the prompt to a representation space.
2. Next, a model called the **prior** maps the text encoding to a corresponding **image** **encoding** that captures the semantic information of the prompt contained in the text encoding.
3. Finally, an **image decoder** stochastically generates an image which is a visual manifestation of this semantic information.

From a bird's eye-view, that's all there is to it! Of course, there are plenty of interesting implementation specifics to discuss, which we will get into below.

## How DALL-E 2 Works: A Detailed Look

Now it's time to dive into each of the above steps separately. Let's get started by looking at how DALL-E 2 learns to link related textual and visual abstractions.

### Step 1 - Linking Textual and Visual Semantics

After inputting **"a teddy bear riding a skateboard in Times Square"**, DALL-E 2 outputs the corresponding image.

How does DALL-E 2 know how a textual concept like "teddy bear" is manifested in the visual space? The **link between textual semantics and their visual representations** in DALL-E 2 is learned by another OpenAI model called **CLIP**( **C** ontrastive **L** anguage- **I** mage **P** re-training).

CLIP is trained on hundreds of millions of images and their associated captions, learning _how much_ a given text snippet relates to an image. This **contrastive** rather than **predictive** objective allows CLIP to learn the link between textual and visual representations of the same abstract object. The entire DALL-E 2 model hinges on CLIP's ability to learn semantics from natural language.

#### CLIP Training

The fundamental principles of training CLIP are quite simple:

1. First, all images and their associated captions are passed through their respective encoders, mapping all objects into an _m-_ dimensional space.
2. Then, the cosine similarity of each _(image, text)_ pair is computed.
3. The training objective is to simultaneously **maximize the cosine similarity** between N **correct** encoded image/caption pairs and **minimize the cosine similarity** between N2 - N **incorrect** encoded image/caption pairs.

This training process is visualized below:

#### Significance of CLIP to DALL-E 2

CLIP is important to DALL-E 2 because **it is what ultimately determines how semantically-related** a natural language snippet is to a visual concept, which is critical for _text-conditional_ image generation.

### Step 2 - Generating Images from Visual Semantics

After training, the CLIP model is frozen and DALL-E 2 moves onto its next task - learning to _reverse_ the image encoding mapping that CLIP just learned. In particular, OpenAI employs a modified version of another one of its previous models, [**GLIDE**](https://arxiv.org/abs/2112.10741?ref=assemblyai.com), to perform this image generation. The GLIDE model learns to _invert_ the image encoding process in order to stochastically decode CLIP image embeddings.

#### What is a Diffusion Model?

Diffusion Models are a thermodynamics-inspired invention that have significantly grown in popularity in recent years. Diffusion Models learn to generate data by _reversing a gradual noising process_. The Diffusion Model learns to navigate backwards along this chain, gradually removing the noise over a series of timesteps to reverse this process.

#### GLIDE Training

While GLIDE was not the first Diffusion Model, its important contribution was in modifying them to allow for **text-conditional image generation**. GLIDE extends the core concept of Diffusion Models by **augmenting the training process with additional textual information**, ultimately resulting in text-conditional image generation. Let's take a look at the training process for GLIDE.

#### Significance of GLIDE to DALL-E 2

GLIDE is important to DALL-E 2 because it allowed the authors to easily port over GLIDE's text-conditional photorealistic image generation capabilities to DALL-E 2 by instead conditioning on **image encodings** in the representation space.

### Step 3 - Mapping from Textual Semantics to Corresponding Visual Semantics

Recall that, in addition to our _image_ encoder, CLIP also learns a _text_ encoder. DALL-E 2 uses another model, which the authors call the **prior**, in order to map **from the text encodings** of image captions **to the** **image encodings** of their corresponding images. The DALL-E 2 authors experiment with both Autoregressive Models and Diffusion Models for the prior, but ultimately find that they yield comparable performance. Given that the Diffusion Model is much more computationally efficient, it is selected as the prior for DALL-E 2.

#### Prior Training

The Diffusion Prior in DALL-E 2 consists of a decoder-only Transformer. It operates, with a causal attention mask, on an ordered sequence of the tokenized text/caption, the CLIP text encodings of these tokens, an encoding for the diffusion timestep, and the noised image passed through the CLIP image encoder.

### Step 4 - Putting It All Together

At this point, we have all of DALL-E 2's functional components and need only to chain them together for text-conditional image generation:

1. First the CLIP text encoder maps the image description into the **representation space**.
2. Then the diffusion prior maps from the CLIP text encoding to a **corresponding CLIP image encoding**.
3. Finally, the modified-GLIDE generation model maps from the representation space into the image space via reverse-Diffusion, **generating one of many possible images that conveys the semantic information** within the input caption.

## **What is DALL-E 3?**

[**DALL-E 3**](https://openai.com/dall-e-3?ref=assemblyai.com), announced in September of 2023, is the successor to DALL-E 2. DALL-E 3 promises “significantly more nuance and detail” relative to DALL-E 2 and other previous systems.

In particular, the model seems to have a focus on capturing prompt semantics. DALL-E 3 will be available to ChatGPT plus subscribers and Enterprise customers in October of 2023.

### How does DALL-E 3 work?

DALL-E 3 is currently said to be in a “research preview”, and no accompanying paper or details about how the model works have been released. However, we can surmise partial information about how DALL-E 3 works based on the recent trends in text-to-image models.

DALL-E 3 is tightly integrated into ChatGPT. In particular, OpenAI says that “DALL-E 3 is built **on** ChatGPT” (emphasis added), not **into** ChatGPT - perhaps this phrasing betrays a deeper connection.

### DALL-E 3 vs DALL-E 2

Overall, DALL-E 3 appears to be an improvement over DALL-E 2 along every evaluation axis of note. DALL-E 3’s tight integration with the ChatGPT web UI makes it significantly easier and more intuitive to use, and it will likely see widespread adoption because of this integration.

## Summary

In this article we covered how the world's premier textually-conditioned image generation model works under the hood. DALL-E 2 can generate semantically plausible photorealistic images given a text prompt, can produce images with specific artistic styles, can produce variations of the same salient features represented in different ways, and can modify existing images.

While there is a lot of discussion to be had about DALL-E 2 and its importance to both Deep Learning and the world at large, we draw your attention to 3 **key takeaways** from the development of DALL-E 2

1. First, DALL-E 2 demonstrates the **power of** [**Diffusion Models**](/content/blog/diffusion-models-for-machine-learning-introduction/index.html) in Deep Learning.
2. The second point is to highlight both the need and **power of using natural language as a means to train** **State-of-the-Art Deep Learning models**.
3. Finally, DALL-E 2 **reaffirms the position of Transformers** as supreme for models trained on web-scale datasets given their impressive parallelizability.

## References

1. [**Deep Unsupervised Learning using Nonequilibrium Thermodynamics**](https://arxiv.org/abs/1503.03585?ref=assemblyai.com)  
2. [**Generative Modeling by Estimating Gradients of the Data Distribution**](https://arxiv.org/abs/1907.05600?ref=assemblyai.com)  
3. [**Hierarchical Text-Conditional Image Generation with CLIP Latents**](https://arxiv.org/pdf/2204.06125.pdf?ref=assemblyai.com)  
4. [**Diffusion Models Beat GANs on Image Synthesis**](https://arxiv.org/abs/2105.05233?ref=assemblyai.com)  
5. [**Denoising Diffusion Probabilistic Models**](https://arxiv.org/pdf/2006.11239.pdf?ref=assemblyai.com)  
6. [**Learning Transferable Visual Models From Natural Language Supervision**](https://arxiv.org/pdf/2103.00020.pdf?ref=assemblyai.com)  
7. [**GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models**](https://arxiv.org/pdf/2112.10741.pdf?ref=assemblyai.com)
