Introducing ChatGPT Images 2.0
OpenAI
0:07 Today, we are launching Imagen 2.0.
0:10 If we think of Dali as cave drawings and
0:12 Imagen 1 as ancient art, then Imagen 2.0
0:15 is the Renaissance.
0:17 Imagen 2.0 is the smartest image
0:20 generation model ever built with the
0:21 ability to generate complex, polished,
0:24 and production-ready visuals with
0:25 accurate text and structured design.
0:28 You see, this model isn't just generating
0:31 it's thinking.
0:32 That's right.
0:33 Imagen 2.0 is thinking and
0:35 researching, and it can even search
0:37 [music] the web to generate images with
0:39 the most accurate information available.
0:42 And with that information, [music]
0:43 the model is able to generate
0:44 infographics that explain complex
0:46 systems and images that solve math
0:48 problems with a proof.
0:53 And with new multilingual capabilities,
0:55 you can [music] create visuals with
0:56 multiple languages for the entire world.
0:59 And now, for the first time in image
1:01 generation, you can create multiple
1:03 distinct images at once, so you can
1:04 generate entire magazines with
1:06 structured typography and photorealistic
1:08 photos, full renovation plans for every
1:10 room in your house,
1:12 or manga comics with recurring
1:13 characters and evolving storylines.
1:14 [music] And you can now generate images with 2K
1:18 resolution across multiple aspect ratios
1:21 with extraordinary micro details.
1:25 You see, we are no longer generating
1:26 [music] images to marvel at.
1:28 With Imagen 2.0, we are generating images to
1:30 discover and navigate, to invent and
1:33 build, to dream and explore the world,
1:36 and bring ideas to life.
1:50 A little over a year ago, we launched
1:51 images in ChatGPT.
1:53 People loved it, and
1:54 it was amazing to see the creativity it
1:56 unleashed.
1:57 But today, we're going to blow way past
1:58 that with Imagen 2.0.
2:01 Imagen 2.0 is a huge step forward.
2:02 This is like going from GPT-3 to GPT-5 all at
2:05 once.
2:06 The ability to create incredible
2:08 new images, express creativity, and
2:09 really make beautiful and complex things
2:11 is is quite remarkable.
2:13 I This is an easier thing to just show
2:15 you than talk about, so I'd like to jump
2:16 right in.
2:17 Uh the team really cooked on
2:18 this one, and we can't wait to see what
2:19 you'll do with it.
2:20 Uh it's available right now in ChatGPT and in the API.
2:25 And here's Gabe to tell you more about it.
2:27 So, hey everyone.
2:27 I'm Gabe.
2:28 I'm the research lead for ChatGPT.
2:30 Hi, I'm Kuan.
2:31 I'm Kenji.
2:32 I'm Alex, and we are researchers on the image
2:34 generation team.
2:36 So, I am very excited about this model,
2:39 and um I I think this model is producing
2:42 images with some quality that I think
2:44 it's very hard to explain, but one way I
2:46 would say it is sort of like uh they
2:47 look just normal.
2:49 They just look like
2:49 normal images.
2:50 And uh one experience I had um looking at
2:54 these images is that after you look at
2:56 these enough, you go back and look at
2:57 previous images, and you see all the
2:58 mistakes that the the previous model
3:00 that they didn't even notice before.
3:02 I mean, they looked great at the time, but
3:03 I think these images look so much
3:05 better.
3:06 Um anyway, I'm going to kick off a
3:07 prompt.
3:08 Um Here's a picture of the four of us
3:12 um we took yesterday, and uh we're going
3:14 to try and create a magazine cover from
3:16 this from this from this image.
3:18 So, um one thing about this model is this model
3:22 has a lot of breath and a lot of depth
3:24 to it.
3:24 So, I think it'll be a while, uh
3:26 you know, before everyone discovers all
3:28 the little nooks and crannies of this
3:29 model.
3:30 Uh but one thing we notice about
3:31 this model is that this model is really
3:33 good at design.
3:34 So, you know, it seems to be really really
3:37 deliberate about where it puts the the
3:40 text in the image.
3:42 And um oh, I think This is live stream.
3:45 This is fine.
3:46 Yeah, I think No, I think Maybe I need
3:48 to retry this.
3:50 Okay.
3:50 Yes.
3:52 No, I think it's fine.
3:52 Oh, oh, yeah.
3:54 I think it's fine.
3:54 Okay, everything is good.
3:55 Yes.
3:55 Um okay.
3:56 So, let's take a
3:57 look.
3:59 Um so, it it it's really deliberate
4:02 about where it puts the text, and and
4:03 the design is looks really nice.
4:07 I [sighs] I remember the time where image
4:09 generation could barely generate a
4:11 single word, um you know, without making typos.
4:14 And now, you know, typos are very rare.
4:16 In fact, it's it's very hard to even find a
4:17 single typo.
4:18 That's been one of the
4:19 things that's surprised me about the new
4:20 model is things that I just never
4:22 thought were possible about cohesion and
4:24 not making a typo in complex text and a
4:26 ton of detail in one image.
4:28 It's rare to find a mistake.
4:29 Yes, it's very rare.
4:31 You can even you
4:31 can you can do a whole paragraph or or
4:33 you know, a full page of text without
4:34 making a mistake.
4:35 And uh oh, Or the full
4:37 layout of a magazine.
4:38 Yeah, full layout of magazine.
4:39 So, all the small text seems to be
4:41 really well done, and I I think the
4:42 design is really nice.
4:43 You guys look like a very cool boy band.
4:47 [laughter]
4:47 Yeah.
4:47 Okay.
4:47 So, we'll be releasing two
4:49 uh versions of the model.
4:51 There was the instant version of the model, which is
4:52 what you're seeing here, and there is a
4:54 thinking version of the model.
4:55 So, the thinking version of the model is um
4:58 something that um you can toggle using
5:01 the thinking mode, and it'll be
5:02 available to paid users.
5:04 And what it does is it it deliberates a little bit
5:06 before it actually generates an image.
5:07 So, it's a winds up a really good
5:09 prompt.
5:10 And it can search the web, it
5:11 can do a lot.
5:12 So, I'm going to um
5:14 try out this prompt.
5:15 So, last year, um
5:18 we did a version of this prompt where
5:22 we turned a selfie much more powerful.
5:25 We can actually generate an entire
5:28 uh manga from a single prompt.
5:30 So, we can generate like three pages of a manga
5:33 from a single prompt.
5:34 And uh I'm just
5:35 going to kick off the generation, and
5:36 then uh Kenji will talk a bit more about
5:39 it.
5:39 is you selected thinking mode.
5:40 We're only doing this for paid users for now,
5:42 and it can do a much more complex image.
5:44 That's correct.
5:44 Yeah, so you have to
5:45 select the thinking mode to get to this
5:46 mode, and you can generate multiple
5:48 images at once and a lot of other very
5:50 interesting things Kenji will talk
5:51 about.
5:51 So, I'm going to kick off another
5:52 prompt here.
5:53 Um and I'm not going to spoil it, but it
5:56 has something to do with word duct tape.
5:58 So, yes.
5:59 [gasps and sighs] Uh okay, I'll hand it over to Kuan to
6:01 talk about instant mode.
6:03 Okay, thanks, Gabe.
6:04 Um instant mode is a version
6:06 available to everyone starting from
6:07 today, uh which we think is much better visual
6:11 intelligence compared to our previous
6:13 models.
6:14 Uh I especially want to
6:14 highlight that this is the first image
6:16 model that is actually useful to our
6:19 daily lives.
6:20 Uh as an example, I'm going
6:23 back to the laptop right now, and um
6:26 I'm asking this model for some help to
6:28 buy new clothes for my upcoming summer
6:30 vacation.
6:31 I've heard this launch.
6:32 Uh so, in this prompt, I'm giving it a portrait
6:35 image of me and asking it to suggest me
6:38 like eight different nice summer
6:40 outfits.
6:41 Uh and in this task, the model needs two
6:44 different kinds of visual intelligence.
6:46 One is visual understanding, where it
6:49 actually looks at my image and
6:51 understand how I look like and come up
6:53 with some plans for nice outfits for me.
6:56 And another access is visual generation,
6:59 where it actually turns that uh
7:01 planned layout into coherent and
7:03 organized image.
7:05 Uh and we think we made
7:06 a lot of progress in both of these
7:08 visual understanding and visual
7:10 generation, both of these aspects.
7:12 Uh as a result, being able to handle this kind
7:14 of tasks uh very well.
7:16 Uh and um we now have an output for this, where
7:21 you can see eight different uh real cool
7:24 outfits for me.
7:25 Mhm.
7:26 Kuan, what do you
7:26 like the best?
7:27 Uh I like the first look
7:30 because I prefer something minimal.
7:33 I think it's somehow pretty similar to
7:34 what I'm wearing right now.
7:36 So, I'm Maybe I'm wearing colors.
7:39 Uh Oh.
7:40 So, I'm going to follow up.
7:42 Okay.
7:42 That's fine.
7:43 Yeah.
7:44 I like look for the
7:45 first look.
7:46 I like the first look.
7:48 Mhm.
7:49 Uh can you zoom into it
7:52 and make a same style fashion
7:58 fashion of me one hero
8:03 few alternative views.
8:07 views and detailed clothes.
8:14 Yeah.
8:14 So, here's a prompt.
8:16 I'm going to follow to the model with
8:17 this prompt.
8:19 Um so, so I'm just basically asking it
8:20 to zoom into it and show me how I look I
8:22 will look like when I'm actually with
8:25 this outfit.
8:26 So, while waiting for it,
8:28 I'm going to revisit the first image a
8:29 little bit more.
8:30 Um so, we're back at the laptop.
8:33 Um and one really cool thing about uh this
8:36 image, I think, is that uh all of these
8:39 clothing pieces are labeled with
8:40 corresponding text.
8:42 Like it shows like
8:43 all of these sneakers and fitted tee,
8:46 things like that.
8:47 And all of these are
8:48 looking really like me.
8:50 So, uh this basically shows that our
8:53 model is uh much more capable in in
8:56 delivering a lot of visual figures
8:59 together with a lot of text, uh which
9:01 comes from much improved visual
9:03 intelligence, essentially.
9:05 Uh yes, so now we have the detailed view of myself,
9:08 where you can see me in this outfit, and
9:11 also from like many different uh angles.
9:14 This really like an experience of just
9:16 going to a store and actually trying
9:18 this outfit.
9:19 So, through this demo, I just want to
9:21 highlight that uh this new model is no
9:24 more like an AI image generator that you
9:26 just gives it a prompt and it returns an
9:28 image.
9:28 It's more like an um
9:30 AI that you just interactively talk to.
9:32 And it's just going to respond you using
9:34 images that are very much understandable
9:36 like this.
9:37 Uh now, I'm passing it to Kenji, who
9:38 will be talking about a deeper
9:40 intelligence in our model called
9:41 thinking mode.
9:42 Thanks, Kuan Kuan.
9:44 Uh a major capability that we've
9:45 introduced in this model is the ability
9:47 for image generation to think before it
9:50 produces its final output.
9:51 This is particularly useful for very complex
9:53 prompts, uh for things that require like
9:56 web searches, or require you to output
9:59 multiple images that have to
10:01 uh maintain coherence with each other,
10:02 or even for it to check its work before
10:04 saying, "Hey, here's your final output."
10:06 But, let's just like look over some
10:08 examples of this first.
10:09 Uh Gabe actually kicked off a few of these examples at
10:11 the start of the live stream.
10:13 So, let's go to the one on the phone, which is the
10:15 one of him and Sam, um the selfie
10:17 they've done, and they created a manga
10:19 of it.
10:19 And if we look at the very first
10:22 uh image, we can see, yeah, it does look like Gabe
10:25 and Sam, right?
10:30 But, I think what's even cooler about it
10:31 is that if you look at the follow-up
10:34 images, they still look like Gabe and Sam, and
10:39 they still look like in the style that
10:41 was originally maintained in the first
10:43 uh in the first page.
10:45 And even better is
10:47 that the story should be very consistent
10:48 among pages one, two, and three.
10:53 Now, this is interesting.
10:55 Thanks.
10:55 Uh to see another one of these
10:57 in action, let's look at the other
10:59 example that Gabe kicked off.
11:02 So, to give a little backstory about
11:03 this, uh a few weeks ago, we beta tested
11:06 the instant version of this model on
11:08 Llama Arena under the codename Duct
11:10 Tape.
11:11 A few did like a few of you on the
11:12 internet were like really good
11:13 detectives and deduced that it was us,
11:16 um but we're going to now say it was us.
11:19 And so, in this prompt, we basically
11:21 asked that uh uh basically GPT Images 2 to go and find
11:26 social media reactions to this Duct Tape
11:28 model, and uh basically quote quote
11:31 people.
11:32 Um and so, we see quotes from
11:33 Threads, LinkedIn, Reddit, etc.
11:36 But, I think an even crazier part is that we've
11:38 also asked the model to put a QR code to
11:42 chat.openai.com so that you can try out
11:44 this model right now
11:46 um for yourselves.
11:47 And can we just make sure that it works?
11:50 Yeah, I tried.
11:51 Oh, nice, nice, nice.
11:52 So, uh image generation was thinking allows
11:55 you to do really complex things such as
11:57 So, in this case, web search, synthesize
11:59 answers, and put a QR code all in one
12:02 image.
12:03 But, we have still more, and Alex
12:05 will talk to you about these new
12:06 details.
12:09 Uh So, we've also made a lot of
12:12 improvements in naturalness.
12:15 And let me just kick off a few uh
12:17 prompts.
12:19 So, like uh like Gabe said earlier, our outputs
12:23 can now just look like natural images.
12:27 And uh you can actually trigger this by
12:29 adding something like photorealistic or
12:32 also there are other variations like
12:33 professional photography, and I've shot on iPhone or disposable
12:38 camera.
12:39 Um So, in in like in this first example,
12:43 I'm just pretending we're back in 2015,
12:46 which is when OpenAI was founded, and
12:49 but somehow there's also GPT Images 2.
12:53 [snorts]
12:52 So, and then uh
12:55 This is As you can see, uh
12:57 the model is actually able to replicate
12:59 the tiny imperfections, uh graininess,
13:02 and lighting of the lecture hall.
13:04 Uh Um even all the text on the slide and uh
13:09 lecture plan that model came up with are
13:11 quite coherent.
13:14 And uh beyond these like
13:17 photorealism, I'm also very excited the
13:18 model is much more flexible now, and in
13:21 particular, we can make really wide and
13:23 really tall images uh up to three 1x3 and 3x1.
13:28 So, uh let's take a look at this.
13:30 This is like one of our team's favorite style
13:33 prompts, and it it really demonstrates
13:35 these uh ability to make really tall images, and
13:39 it makes my neck super long.
13:42 And and like I mean, this is pretty cool, but uh
13:47 I mean, maybe it's a bit hard to use as
13:49 a profile picture or share, so you can
13:52 also use this option
13:54 to make a 1x1, and just
13:56 I won't show this for interest of time.
14:00 So, uh For another fun example that combines
14:05 both the aspect ratio and naturalness, I
14:07 I also have this uh
14:09 uh asked the model to make a 360
14:12 uh image of the
14:14 moon landing.
14:16 And I I think it looks like a 360 photo,
14:19 a panorama, but we can also take a look in this panorama
14:24 viewer that I have coded earlier.
14:28 Uh So, Wow.
14:33 As you can see, it's it's a actually a
14:35 very consistent uh 360 image.
14:40 It's You can also see this this the sun, and
14:44 the shadows are are also in the right
14:46 direction.
14:47 Oh, that's super cool.
14:48 That's incredible.
14:49 You said you have I coded
14:50 this part of it?
14:50 Yeah, yeah, just I just
14:51 made it with code extra quickly.
14:54 Um and there's um
14:56 Let's see but you have to look for it.
14:58 Like This is incredible.
15:03 The the images are
15:04 obviously beautiful, but the
15:05 intelligence uh behind these images and
15:07 what a difference that is to any other
15:08 image generation service out there has
15:10 been has been incredible.
15:12 Uh huge congratulations on the progress
15:13 here.
15:14 Uh okay, next up, we're going to
15:16 have Nithanth and Boyuan join us for uh
15:19 a little bit more.
15:21 And while we're doing that, Gabe, I'm
15:22 curious what styles you have been
15:24 enjoying the most or sort of most
15:25 surprised by.
15:26 Yeah, I I think there's a
15:27 few keywords that I really like, but I
15:29 think like Alex said, I think the word
15:31 photorealism actually triggers something
15:33 really very interesting in the model.
15:36 Definitely give that one a
15:37 try.
15:37 Yes.
15:40 Okay, welcome you guys.
15:41 Hello.
15:42 Thank you, Sam, and thank Gabe.
15:44 Um hi, I'm Boyuan.
15:45 I'm another member of
15:46 the image and research team.
15:48 And I'm Nithanth.
15:49 I'm an engineer in the
15:50 ChatGPT Images team.
15:52 I'm about to introduce the improved the
15:54 text text rendering capability of our
15:56 new model.
15:58 OpenAI is a San Francisco-based company.
16:00 We speak English and use English at
16:02 work.
16:03 However, we want everyone in the
16:04 world to enjoy the same excitement we
16:07 have when generating images.
16:09 So, in Imagen 2, we made a lot of
16:12 improvements to make sure that our model
16:14 can generate every text perfectly across
16:17 all the languages, all the cultures in
16:19 the world.
16:20 Let's take a look.
16:22 So, in my first example, I want to
16:25 generate a poster that's a typography
16:27 art about different language in the
16:29 world.
16:29 It's going to feature many many
16:30 languages, and let's see how does it
16:33 appears.
16:34 Um And while it's generating, I'm going
16:37 to kick off another demo.
16:38 Let's say I want to open OpenAI Bakery.
16:41 It's a fictional bakery, and I want to open it
16:43 in Japan.
16:44 I want to make a poster about
16:46 it and put it in Japanese.
16:49 What languages have you noticed that the
16:51 new model's gotten the best at?
16:53 Um I think mostly Asian languages.
16:56 Let's say Hindi, Chinese, Korean, and um
16:59 Japanese.
17:00 That's because those languages
17:01 traditionally have thousands of
17:03 characters in the alphabet, unlike the
17:05 26 in English.
17:07 So, um previously, our
17:09 model had a hard time memorizing these
17:11 characters, but now just prompt it and
17:14 and generate entire pages of text in
17:16 these languages without errors.
17:18 Wow.
17:19 Nice.
17:21 Let's see how does it go.
17:22 Oh, here is our first example, the typography art.
17:25 I deliberately to be in the form of
17:28 photography of a uh of a real magazine.
17:31 Hm.
17:31 So, it not only look realistic, but I can also see
17:34 the correct characters.
17:35 Here is "ni hao"
17:36 in Chinese.
17:37 There is "hello" as well,
17:38 "bonjour" in French.
17:39 And I hope everyone
17:40 in the world can actually enjoy our
17:41 model creating your own art using your
17:43 own language.
17:44 Let's take a look at the second example,
17:46 my OpenAI Bakery.
17:47 Oh, look at it.
17:48 It even made our logo into
17:50 into this piece of bread, right?
17:52 Um this is a Japanese poster.
17:54 You can see all
17:55 the kanji, all the hiragana.
17:57 Uh you can even zoom in and see the details.
17:59 Look at this.
18:01 Look at all the hiraganas here.
18:03 Hm.
18:04 So, I really hope everyone in the world can
18:07 use this model to make your own poster,
18:08 open your own shop, everything.
18:10 And just to show everyone how far we can go with
18:14 our image generation model.
18:16 So, this is an image I generated with our
18:19 experimental 4K API.
18:21 Um this is just a
18:22 pile of rice, but this is also not just
18:24 one pile of rice.
18:25 What if I tell you
18:26 there's one single grain in it with the
18:29 text GPT Image on it?
18:31 Can you find it?
18:32 Here we go.
18:34 Yeah, this is the same I made
18:36 I made it easy for you guys.
18:38 some text.
18:39 zoom in.
18:40 GPT Image 2 on one single grain
18:42 of rice among the entire
18:46 Hm.
18:45 pile this big.
18:47 This is how far we can go
18:48 with our latest model.
18:49 Amazing.
18:50 Next, I will let Nithanth take over.
18:53 Yeah, so Imagen 2.0 is available to all
18:56 users to try out right now.
18:58 And if you're accessing ChatGPT from your app,
19:00 um make sure to update it to the latest
19:01 version, and you should see a welcome
19:03 screen that looks like this, which means
19:04 you're good to go.
19:06 Uh I'm going to start off with a a
19:08 simple everyday kind of prompt.
19:10 I'm asking it to um create a recipe in
19:13 Hindi.
19:14 Um as Boyuan said, the the new
19:16 model is significantly better at
19:18 understanding and rendering text in lots
19:20 of languages, including many Indian ones
19:23 that I've tried like Hindi, Telugu,
19:25 Kannada, Tamil, um Marathi, and so on.
19:29 And the difference is especially obvious
19:31 if uh there's a lot of densely packed
19:33 text.
19:34 So, let's see what it comes back
19:35 with.
19:36 I'm also curious to see what Indian dish
19:38 it decides to go with.
19:45 Up, there we go.
19:47 It went with aloo paratha.
19:48 That's a that's a classic.
19:50 Oh.
19:52 And uh the text looks really good, too.
19:56 I don't spot any any errors
19:58 um at first glance.
20:02 And next, let's also check out some of the
20:05 uh the new preset styles that we've
20:07 added into the app.
20:08 Just going to select create images here.
20:11 And you'll see a bunch of fun ones and
20:13 some that really take advantage of the
20:15 new models' capabilities.
20:17 Um actually, how about why don't we make
20:20 um logos for the OpenAI bakery, Bojan?
20:23 Sure.
20:24 Why don't you just take a photo of
20:26 my bakery poster and see what fixing.
20:29 Let's do it.
20:33 So, looks like this is going to come
20:34 back with 16 to 20 um logo ideas.
20:38 Uh but this is actually a rather uh simple
20:41 prompt uh given the model's
20:42 capabilities.
20:43 It's really uh good at
20:45 following very detailed instructions.
20:47 So, uh if you have very specific brand
20:49 language, design aesthetics, um
20:52 all of those things that really matter
20:53 for creative work, um you can use
20:55 ChatGPT to iterate and refine on your
20:58 ideas to get exactly what you want out
21:00 of it.
21:03 And we have colorful logo ideas right here.
21:07 Wow, here we go.
21:08 Nice.
21:09 Which one do you
21:09 guys like the most?
21:11 These are good.
21:13 Uh how about this one?
21:14 Oh, yeah.
21:15 This one combines the our logo and the bread.
21:18 I like it.
21:18 All of this is also uh making
21:20 me hungry.
21:23 This was uh This is really amazing.
21:25 I can't wait to
21:26 see what people will do with this.
21:27 The The beauty of the images uh will come
21:30 through right away.
21:31 The intelligence is very deep, and we hope you all have fun
21:33 exploring this.
21:34 As we mentioned, this is
21:35 live today in ChatGPT and the API.
21:38 Uh so proud of the team on what they've
21:40 created here, and we hope you all have
21:42 as much fun using it as we did getting
21:44 to build it.
21:44 Thank you very much.