What is a Multi-Modal Model?
Per Wikipedia - a multi-modal model integrates and processes multiple types of data, such as text and images. This integration allows for a more holistic understanding of how to integrate start images as part of text to image generation, along with the ability to be prompted to "think" and use those thoughts as part of the response - in the case on NightCafe, a created image.
This collection is a series on how you can use these in ways different to the prompting you may have done before with other models. Right now the two multi-modal models on the site are Image GPT, and Gemini Flash 2.5.
To start off, a brief demonstration within this creation of what Gemini is capable of. The prompt used is like a prompt you could submit to a large language model, such as ChatGPT. But in this case, we asked for answers and then provided instructions to use those answers in the generated image.