Introducing ChatGPT Images 2.0

Introducing ChatGPT Images 2.0

OpenAI

0:07 Today, we are launching Imagen 2.0.

0:10 If we think of Dali as cave drawings and

0:12 Imagen 1 as ancient art, then Imagen 2.0

0:15 is the Renaissance.

0:17 Imagen 2.0 is the smartest image

0:20 generation model ever built with the

0:21 ability to generate complex, polished,

0:24 and production-ready visuals with

0:25 accurate text and structured design.

0:28 You see, this model isn't just generating

0:31 it's thinking.

0:32 That's right.

0:33 Imagen 2.0 is thinking and

0:35 researching, and it can even search

0:37 [music] the web to generate images with

0:39 the most accurate information available.

0:42 And with that information, [music]

0:43 the model is able to generate

0:44 infographics that explain complex

0:46 systems and images that solve math

0:48 problems with a proof.

0:53 And with new multilingual capabilities,

0:55 you can [music] create visuals with

0:56 multiple languages for the entire world.

0:59 And now, for the first time in image

1:01 generation, you can create multiple

1:03 distinct images at once, so you can

1:04 generate entire magazines with

1:06 structured typography and photorealistic

1:08 photos, full renovation plans for every

1:10 room in your house,

1:12 or manga comics with recurring

1:13 characters and evolving storylines.

1:14 [music] And you can now generate images with 2K

1:18 resolution across multiple aspect ratios

1:21 with extraordinary micro details.

1:25 You see, we are no longer generating

1:26 [music] images to marvel at.

1:28 With Imagen 2.0, we are generating images to

1:30 discover and navigate, to invent and

1:33 build, to dream and explore the world,

1:36 and bring ideas to life.

1:50 A little over a year ago, we launched

1:51 images in ChatGPT.

1:53 People loved it, and

1:54 it was amazing to see the creativity it

1:56 unleashed.

1:57 But today, we're going to blow way past

1:58 that with Imagen 2.0.

2:01 Imagen 2.0 is a huge step forward.

2:02 This is like going from GPT-3 to GPT-5 all at

2:05 once.

2:06 The ability to create incredible

2:08 new images, express creativity, and

2:09 really make beautiful and complex things

2:11 is is quite remarkable.

2:13 I This is an easier thing to just show

2:15 you than talk about, so I'd like to jump

2:16 right in.

2:17 Uh the team really cooked on

2:18 this one, and we can't wait to see what

2:19 you'll do with it.

2:20 Uh it's available right now in ChatGPT and in the API.

2:25 And here's Gabe to tell you more about it.

2:27 So, hey everyone.

2:27 I'm Gabe.

2:28 I'm the research lead for ChatGPT.

2:30 Hi, I'm Kuan.

2:31 I'm Kenji.

2:32 I'm Alex, and we are researchers on the image

2:34 generation team.

2:36 So, I am very excited about this model,

2:39 and um I I think this model is producing

2:42 images with some quality that I think

2:44 it's very hard to explain, but one way I

2:46 would say it is sort of like uh they

2:47 look just normal.

2:49 They just look like

2:49 normal images.

2:50 And uh one experience I had um looking at

2:54 these images is that after you look at

2:56 these enough, you go back and look at

2:57 previous images, and you see all the

2:58 mistakes that the the previous model

3:00 that they didn't even notice before.

3:02 I mean, they looked great at the time, but

3:03 I think these images look so much

3:05 better.

3:06 Um anyway, I'm going to kick off a

3:07 prompt.

3:08 Um Here's a picture of the four of us

3:12 um we took yesterday, and uh we're going

3:14 to try and create a magazine cover from

3:16 this from this from this image.

3:18 So, um one thing about this model is this model

3:22 has a lot of breath and a lot of depth

3:24 to it.

3:24 So, I think it'll be a while, uh

3:26 you know, before everyone discovers all

3:28 the little nooks and crannies of this

3:29 model.

3:30 Uh but one thing we notice about

3:31 this model is that this model is really

3:33 good at design.

3:34 So, you know, it seems to be really really

3:37 deliberate about where it puts the the

3:40 text in the image.

3:42 And um oh, I think This is live stream.

3:45 This is fine.

3:46 Yeah, I think No, I think Maybe I need

3:48 to retry this.

3:50 Okay.

3:50 Yes.

3:52 No, I think it's fine.

3:52 Oh, oh, yeah.

3:54 I think it's fine.

3:54 Okay, everything is good.

3:55 Yes.

3:55 Um okay.

3:56 So, let's take a

3:57 look.

3:59 Um so, it it it's really deliberate

4:02 about where it puts the text, and and

4:03 the design is looks really nice.

4:07 I [sighs] I remember the time where image

4:09 generation could barely generate a

4:11 single word, um you know, without making typos.

4:14 And now, you know, typos are very rare.

4:16 In fact, it's it's very hard to even find a

4:17 single typo.

4:18 That's been one of the

4:19 things that's surprised me about the new

4:20 model is things that I just never

4:22 thought were possible about cohesion and

4:24 not making a typo in complex text and a

4:26 ton of detail in one image.

4:28 It's rare to find a mistake.

4:29 Yes, it's very rare.

4:31 You can even you

4:31 can you can do a whole paragraph or or

4:33 you know, a full page of text without

4:34 making a mistake.

4:35 And uh oh, Or the full

4:37 layout of a magazine.

4:38 Yeah, full layout of magazine.

4:39 So, all the small text seems to be

4:41 really well done, and I I think the

4:42 design is really nice.

4:43 You guys look like a very cool boy band.

4:47 [laughter]

4:47 Yeah.

4:47 Okay.

4:47 So, we'll be releasing two

4:49 uh versions of the model.

4:51 There was the instant version of the model, which is

4:52 what you're seeing here, and there is a

4:54 thinking version of the model.

4:55 So, the thinking version of the model is um

4:58 something that um you can toggle using

5:01 the thinking mode, and it'll be

5:02 available to paid users.

5:04 And what it does is it it deliberates a little bit

5:06 before it actually generates an image.

5:07 So, it's a winds up a really good

5:09 prompt.

5:10 And it can search the web, it

5:11 can do a lot.

5:12 So, I'm going to um

5:14 try out this prompt.

5:15 So, last year, um

5:18 we did a version of this prompt where

5:22 we turned a selfie much more powerful.

5:25 We can actually generate an entire

5:28 uh manga from a single prompt.

5:30 So, we can generate like three pages of a manga

5:33 from a single prompt.

5:34 And uh I'm just

5:35 going to kick off the generation, and

5:36 then uh Kenji will talk a bit more about

5:39 it.

5:39 is you selected thinking mode.

5:40 We're only doing this for paid users for now,

5:42 and it can do a much more complex image.

5:44 That's correct.

5:44 Yeah, so you have to

5:45 select the thinking mode to get to this

5:46 mode, and you can generate multiple

5:48 images at once and a lot of other very

5:50 interesting things Kenji will talk

5:51 about.

5:51 So, I'm going to kick off another

5:52 prompt here.

5:53 Um and I'm not going to spoil it, but it

5:56 has something to do with word duct tape.

5:58 So, yes.

5:59 [gasps and sighs] Uh okay, I'll hand it over to Kuan to

6:01 talk about instant mode.

6:03 Okay, thanks, Gabe.

6:04 Um instant mode is a version

6:06 available to everyone starting from

6:07 today, uh which we think is much better visual

6:11 intelligence compared to our previous

6:13 models.

6:14 Uh I especially want to

6:14 highlight that this is the first image

6:16 model that is actually useful to our

6:19 daily lives.

6:20 Uh as an example, I'm going

6:23 back to the laptop right now, and um

6:26 I'm asking this model for some help to

6:28 buy new clothes for my upcoming summer

6:30 vacation.

6:31 I've heard this launch.

6:32 Uh so, in this prompt, I'm giving it a portrait

6:35 image of me and asking it to suggest me

6:38 like eight different nice summer

6:40 outfits.

6:41 Uh and in this task, the model needs two

6:44 different kinds of visual intelligence.

6:46 One is visual understanding, where it

6:49 actually looks at my image and

6:51 understand how I look like and come up

6:53 with some plans for nice outfits for me.

6:56 And another access is visual generation,

6:59 where it actually turns that uh

7:01 planned layout into coherent and

7:03 organized image.

7:05 Uh and we think we made

7:06 a lot of progress in both of these

7:08 visual understanding and visual

7:10 generation, both of these aspects.

7:12 Uh as a result, being able to handle this kind

7:14 of tasks uh very well.

7:16 Uh and um we now have an output for this, where

7:21 you can see eight different uh real cool

7:24 outfits for me.

7:25 Mhm.

7:26 Kuan, what do you

7:26 like the best?

7:27 Uh I like the first look

7:30 because I prefer something minimal.

7:33 I think it's somehow pretty similar to

7:34 what I'm wearing right now.

7:36 So, I'm Maybe I'm wearing colors.

7:39 Uh Oh.

7:40 So, I'm going to follow up.

7:42 Okay.

7:42 That's fine.

7:43 Yeah.

7:44 I like look for the

7:45 first look.

7:46 I like the first look.

7:48 Mhm.

7:49 Uh can you zoom into it

7:52 and make a same style fashion

7:58 fashion of me one hero

8:03 few alternative views.

8:07 views and detailed clothes.

8:14 Yeah.

8:14 So, here's a prompt.

8:16 I'm going to follow to the model with

8:17 this prompt.

8:19 Um so, so I'm just basically asking it

8:20 to zoom into it and show me how I look I

8:22 will look like when I'm actually with

8:25 this outfit.

8:26 So, while waiting for it,

8:28 I'm going to revisit the first image a

8:29 little bit more.

8:30 Um so, we're back at the laptop.

8:33 Um and one really cool thing about uh this

8:36 image, I think, is that uh all of these

8:39 clothing pieces are labeled with

8:40 corresponding text.

8:42 Like it shows like

8:43 all of these sneakers and fitted tee,

8:46 things like that.

8:47 And all of these are

8:48 looking really like me.

8:50 So, uh this basically shows that our

8:53 model is uh much more capable in in

8:56 delivering a lot of visual figures

8:59 together with a lot of text, uh which

9:01 comes from much improved visual

9:03 intelligence, essentially.

9:05 Uh yes, so now we have the detailed view of myself,

9:08 where you can see me in this outfit, and

9:11 also from like many different uh angles.

9:14 This really like an experience of just

9:16 going to a store and actually trying

9:18 this outfit.

9:19 So, through this demo, I just want to

9:21 highlight that uh this new model is no

9:24 more like an AI image generator that you

9:26 just gives it a prompt and it returns an

9:28 image.

9:28 It's more like an um

9:30 AI that you just interactively talk to.

9:32 And it's just going to respond you using

9:34 images that are very much understandable

9:36 like this.

9:37 Uh now, I'm passing it to Kenji, who

9:38 will be talking about a deeper

9:40 intelligence in our model called

9:41 thinking mode.

9:42 Thanks, Kuan Kuan.

9:44 Uh a major capability that we've

9:45 introduced in this model is the ability

9:47 for image generation to think before it

9:50 produces its final output.

9:51 This is particularly useful for very complex

9:53 prompts, uh for things that require like

9:56 web searches, or require you to output

9:59 multiple images that have to

10:01 uh maintain coherence with each other,

10:02 or even for it to check its work before

10:04 saying, "Hey, here's your final output."

10:06 But, let's just like look over some

10:08 examples of this first.

10:09 Uh Gabe actually kicked off a few of these examples at

10:11 the start of the live stream.

10:13 So, let's go to the one on the phone, which is the

10:15 one of him and Sam, um the selfie

10:17 they've done, and they created a manga

10:19 of it.

10:19 And if we look at the very first

10:22 uh image, we can see, yeah, it does look like Gabe

10:25 and Sam, right?

10:30 But, I think what's even cooler about it

10:31 is that if you look at the follow-up

10:34 images, they still look like Gabe and Sam, and

10:39 they still look like in the style that

10:41 was originally maintained in the first

10:43 uh in the first page.

10:45 And even better is

10:47 that the story should be very consistent

10:48 among pages one, two, and three.

10:53 Now, this is interesting.

10:55 Thanks.

10:55 Uh to see another one of these

10:57 in action, let's look at the other

10:59 example that Gabe kicked off.

11:02 So, to give a little backstory about

11:03 this, uh a few weeks ago, we beta tested

11:06 the instant version of this model on

11:08 Llama Arena under the codename Duct

11:10 Tape.

11:11 A few did like a few of you on the

11:12 internet were like really good

11:13 detectives and deduced that it was us,

11:16 um but we're going to now say it was us.

11:19 And so, in this prompt, we basically

11:21 asked that uh uh basically GPT Images 2 to go and find

11:26 social media reactions to this Duct Tape

11:28 model, and uh basically quote quote

11:31 people.

11:32 Um and so, we see quotes from

11:33 Threads, LinkedIn, Reddit, etc.

11:36 But, I think an even crazier part is that we've

11:38 also asked the model to put a QR code to

11:42 chat.openai.com so that you can try out

11:44 this model right now

11:46 um for yourselves.

11:47 And can we just make sure that it works?

11:50 Yeah, I tried.

11:51 Oh, nice, nice, nice.

11:52 So, uh image generation was thinking allows

11:55 you to do really complex things such as

11:57 So, in this case, web search, synthesize

11:59 answers, and put a QR code all in one

12:02 image.

12:03 But, we have still more, and Alex

12:05 will talk to you about these new

12:06 details.

12:09 Uh So, we've also made a lot of

12:12 improvements in naturalness.

12:15 And let me just kick off a few uh

12:17 prompts.

12:19 So, like uh like Gabe said earlier, our outputs

12:23 can now just look like natural images.

12:27 And uh you can actually trigger this by

12:29 adding something like photorealistic or

12:32 also there are other variations like

12:33 professional photography, and I've shot on iPhone or disposable

12:38 camera.

12:39 Um So, in in like in this first example,

12:43 I'm just pretending we're back in 2015,

12:46 which is when OpenAI was founded, and

12:49 but somehow there's also GPT Images 2.

12:53 [snorts]

12:52 So, and then uh

12:55 This is As you can see, uh

12:57 the model is actually able to replicate

12:59 the tiny imperfections, uh graininess,

13:02 and lighting of the lecture hall.

13:04 Uh Um even all the text on the slide and uh

13:09 lecture plan that model came up with are

13:11 quite coherent.

13:14 And uh beyond these like

13:17 photorealism, I'm also very excited the

13:18 model is much more flexible now, and in

13:21 particular, we can make really wide and

13:23 really tall images uh up to three 1x3 and 3x1.

13:28 So, uh let's take a look at this.

13:30 This is like one of our team's favorite style

13:33 prompts, and it it really demonstrates

13:35 these uh ability to make really tall images, and

13:39 it makes my neck super long.

13:42 And and like I mean, this is pretty cool, but uh

13:47 I mean, maybe it's a bit hard to use as

13:49 a profile picture or share, so you can

13:52 also use this option

13:54 to make a 1x1, and just

13:56 I won't show this for interest of time.

14:00 So, uh For another fun example that combines

14:05 both the aspect ratio and naturalness, I

14:07 I also have this uh

14:09 uh asked the model to make a 360

14:12 uh image of the

14:14 moon landing.

14:16 And I I think it looks like a 360 photo,

14:19 a panorama, but we can also take a look in this panorama

14:24 viewer that I have coded earlier.

14:28 Uh So, Wow.

14:33 As you can see, it's it's a actually a

14:35 very consistent uh 360 image.

14:40 It's You can also see this this the sun, and

14:44 the shadows are are also in the right

14:46 direction.

14:47 Oh, that's super cool.

14:48 That's incredible.

14:49 You said you have I coded

14:50 this part of it?

14:50 Yeah, yeah, just I just

14:51 made it with code extra quickly.

14:54 Um and there's um

14:56 Let's see but you have to look for it.

14:58 Like This is incredible.

15:03 The the images are

15:04 obviously beautiful, but the

15:05 intelligence uh behind these images and

15:07 what a difference that is to any other

15:08 image generation service out there has

15:10 been has been incredible.

15:12 Uh huge congratulations on the progress

15:13 here.

15:14 Uh okay, next up, we're going to

15:16 have Nithanth and Boyuan join us for uh

15:19 a little bit more.

15:21 And while we're doing that, Gabe, I'm

15:22 curious what styles you have been

15:24 enjoying the most or sort of most

15:25 surprised by.

15:26 Yeah, I I think there's a

15:27 few keywords that I really like, but I

15:29 think like Alex said, I think the word

15:31 photorealism actually triggers something

15:33 really very interesting in the model.

15:36 Definitely give that one a

15:37 try.

15:37 Yes.

15:40 Okay, welcome you guys.

15:41 Hello.

15:42 Thank you, Sam, and thank Gabe.

15:44 Um hi, I'm Boyuan.

15:45 I'm another member of

15:46 the image and research team.

15:48 And I'm Nithanth.

15:49 I'm an engineer in the

15:50 ChatGPT Images team.

15:52 I'm about to introduce the improved the

15:54 text text rendering capability of our

15:56 new model.

15:58 OpenAI is a San Francisco-based company.

16:00 We speak English and use English at

16:02 work.

16:03 However, we want everyone in the

16:04 world to enjoy the same excitement we

16:07 have when generating images.

16:09 So, in Imagen 2, we made a lot of

16:12 improvements to make sure that our model

16:14 can generate every text perfectly across

16:17 all the languages, all the cultures in

16:19 the world.

16:20 Let's take a look.

16:22 So, in my first example, I want to

16:25 generate a poster that's a typography

16:27 art about different language in the

16:29 world.

16:29 It's going to feature many many

16:30 languages, and let's see how does it

16:33 appears.

16:34 Um And while it's generating, I'm going

16:37 to kick off another demo.

16:38 Let's say I want to open OpenAI Bakery.

16:41 It's a fictional bakery, and I want to open it

16:43 in Japan.

16:44 I want to make a poster about

16:46 it and put it in Japanese.

16:49 What languages have you noticed that the

16:51 new model's gotten the best at?

16:53 Um I think mostly Asian languages.

16:56 Let's say Hindi, Chinese, Korean, and um

16:59 Japanese.

17:00 That's because those languages

17:01 traditionally have thousands of

17:03 characters in the alphabet, unlike the

17:05 26 in English.

17:07 So, um previously, our

17:09 model had a hard time memorizing these

17:11 characters, but now just prompt it and

17:14 and generate entire pages of text in

17:16 these languages without errors.

17:18 Wow.

17:19 Nice.

17:21 Let's see how does it go.

17:22 Oh, here is our first example, the typography art.

17:25 I deliberately to be in the form of

17:28 photography of a uh of a real magazine.

17:31 Hm.

17:31 So, it not only look realistic, but I can also see

17:34 the correct characters.

17:35 Here is "ni hao"

17:36 in Chinese.

17:37 There is "hello" as well,

17:38 "bonjour" in French.

17:39 And I hope everyone

17:40 in the world can actually enjoy our

17:41 model creating your own art using your

17:43 own language.

17:44 Let's take a look at the second example,

17:46 my OpenAI Bakery.

17:47 Oh, look at it.

17:48 It even made our logo into

17:50 into this piece of bread, right?

17:52 Um this is a Japanese poster.

17:54 You can see all

17:55 the kanji, all the hiragana.

17:57 Uh you can even zoom in and see the details.

17:59 Look at this.

18:01 Look at all the hiraganas here.

18:03 Hm.

18:04 So, I really hope everyone in the world can

18:07 use this model to make your own poster,

18:08 open your own shop, everything.

18:10 And just to show everyone how far we can go with

18:14 our image generation model.

18:16 So, this is an image I generated with our

18:19 experimental 4K API.

18:21 Um this is just a

18:22 pile of rice, but this is also not just

18:24 one pile of rice.

18:25 What if I tell you

18:26 there's one single grain in it with the

18:29 text GPT Image on it?

18:31 Can you find it?

18:32 Here we go.

18:34 Yeah, this is the same I made

18:36 I made it easy for you guys.

18:38 some text.

18:39 zoom in.

18:40 GPT Image 2 on one single grain

18:42 of rice among the entire

18:46 Hm.

18:45 pile this big.

18:47 This is how far we can go

18:48 with our latest model.

18:49 Amazing.

18:50 Next, I will let Nithanth take over.

18:53 Yeah, so Imagen 2.0 is available to all

18:56 users to try out right now.

18:58 And if you're accessing ChatGPT from your app,

19:00 um make sure to update it to the latest

19:01 version, and you should see a welcome

19:03 screen that looks like this, which means

19:04 you're good to go.

19:06 Uh I'm going to start off with a a

19:08 simple everyday kind of prompt.

19:10 I'm asking it to um create a recipe in

19:13 Hindi.

19:14 Um as Boyuan said, the the new

19:16 model is significantly better at

19:18 understanding and rendering text in lots

19:20 of languages, including many Indian ones

19:23 that I've tried like Hindi, Telugu,

19:25 Kannada, Tamil, um Marathi, and so on.

19:29 And the difference is especially obvious

19:31 if uh there's a lot of densely packed

19:33 text.

19:34 So, let's see what it comes back

19:35 with.

19:36 I'm also curious to see what Indian dish

19:38 it decides to go with.

19:45 Up, there we go.

19:47 It went with aloo paratha.

19:48 That's a that's a classic.

19:50 Oh.

19:52 And uh the text looks really good, too.

19:56 I don't spot any any errors

19:58 um at first glance.

20:02 And next, let's also check out some of the

20:05 uh the new preset styles that we've

20:07 added into the app.

20:08 Just going to select create images here.

20:11 And you'll see a bunch of fun ones and

20:13 some that really take advantage of the

20:15 new models' capabilities.

20:17 Um actually, how about why don't we make

20:20 um logos for the OpenAI bakery, Bojan?

20:23 Sure.

20:24 Why don't you just take a photo of

20:26 my bakery poster and see what fixing.

20:29 Let's do it.

20:33 So, looks like this is going to come

20:34 back with 16 to 20 um logo ideas.

20:38 Uh but this is actually a rather uh simple

20:41 prompt uh given the model's

20:42 capabilities.

20:43 It's really uh good at

20:45 following very detailed instructions.

20:47 So, uh if you have very specific brand

20:49 language, design aesthetics, um

20:52 all of those things that really matter

20:53 for creative work, um you can use

20:55 ChatGPT to iterate and refine on your

20:58 ideas to get exactly what you want out

21:00 of it.

21:03 And we have colorful logo ideas right here.

21:07 Wow, here we go.

21:08 Nice.

21:09 Which one do you

21:09 guys like the most?

21:11 These are good.

21:13 Uh how about this one?

21:14 Oh, yeah.

21:15 This one combines the our logo and the bread.

21:18 I like it.

21:18 All of this is also uh making

21:20 me hungry.

21:23 This was uh This is really amazing.

21:25 I can't wait to

21:26 see what people will do with this.

21:27 The The beauty of the images uh will come

21:30 through right away.

21:31 The intelligence is very deep, and we hope you all have fun

21:33 exploring this.

21:34 As we mentioned, this is

21:35 live today in ChatGPT and the API.

21:38 Uh so proud of the team on what they've

21:40 created here, and we hope you all have

21:42 as much fun using it as we did getting

21:44 to build it.

21:44 Thank you very much.

Study with Looplines Download Captions Watch on YouTube