For over a year, the term natively multimodal has resonated in the world of artificial intelligence, but few have managed to fully leverage these capabilities. Now, Google has made its move with the launch of its latest model, Gemini 2.0 Flash Experimental, which allows not only generating images but also editing them natively. In other words, it has given Photoshop a pinch 🤏🏼…
Why is image generation so important? Although AI image generation has been available through chatbots like ChatGPT, these often rely on specialized models like Dall-E 3 or Imagen 3, which are extensions of the main model and not an integral part of it. In contrast, models like Gemini are natively multimodal, meaning they can understand and create both text and images intrinsically.
Natively Image Generation with Gemini 2.0 Flash Experimental
Currently, this native image generation feature is not available to all users. The Gemini 2.0 Flash Experimental model can be tested for free at the Google AI Studio and will soon be available to a wider audience. After experimenting with this model, I can say that the experience was truly remarkable.
I started by asking Gemini to create a visual guide on how to make Bolognese macaroni. The results were astonishing, showing a remarkable consistency among the generated images, from the pan to the ingredients. Each image maintains the same resolution of 1024 x 680, making it easy to create visual guides on any topic.

Then, I asked Gemini to generate an empty room, and I kept asking for modifications on the decoration and utility of the room. The continuity it maintained was astonishing.

Natively Image Editing with Gemini 2.0 Flash Experimental
To demonstrate the image editing feature, I uploaded a photo of my garage and asked it to change my car to a white Tesla, and the result was impressive. Finally, I asked it to add some tables with computers, and this showed me the potential of image editing thanks to Gemini’s native multimodal capability. They weren’t perfect, but they were very good. Additionally, I asked Gemini to colorize an old black and white photo, and the result exceeded my expectations, with optimal visual quality and no visible errors.
The possibilities with Gemini are vast and exciting. Google has done an admirable job of integrating image generation and editing natively. With the recent launch of Veo 2 for video generation and Imagen 3 for specialized image generation, it seems that Google has surpassed OpenAI in several aspects, not just in text generation. It will be interesting to see how OpenAI responds to this advancement with its ChatGPT.












0 Comments