ChatGPT is made from 100 million of these [The Perceptron]

ChatGPT is made from 100 million of these [The Perceptron]

Welch Labs

0:00 this is a perceptron the machine shocked

0:02 the world in the 1950s by learning to recognize

0:05 patterns completely automatically and today the algorithm it

0:09 implements has become the core building block of AI

0:11 systems like chat GPT but why is

0:14 this the atomic unit of the intelligent systems we

0:16 have today the perceptron works by processing patterns

0:20 we input using these switches like this t-shape

0:23 switches in the up position output a positive

0:25 voltage and switches in the down position output

0:27 a negative voltage each switch is connected to an indicator

0:30 LED and then to one of these dials

0:33 rotating a dial multiplies the output of the switch

0:36 by the number shown on the dial and the meter shows the result of adding

0:40 the signals from all the dials together now

0:43 is there a way to configure our dials

0:45 such that the perceptron always outputs a positive

0:48 signal for t-shapes while outputting a negative value

0:51 for other types of shapes like this J shape note that our shapes won't be

0:55 in the same position each time remarkably it turns

0:58 out that if a configuration of that solves

1:00 our problem exists there is a simple procedure we

1:03 can follow that is guaranteed to find it

1:05 every time starting with this t-shape our meter

1:08 is showing a value close to zero but we want it to be positive in this case

1:13 our procedure tells us to turn all the knobs that are switch on to the right

1:17 by a constant value called The Learning rate

1:19 and to turn all the knobs that are Switched

1:21 Off to the left by the Learning rate our next pattern is a j-shape so we

1:26 want our machine to Output a negative value

1:28 however the current dial configuration outputs a positive

1:31 value in this case our procedure tells us to turn down all the dials that are

1:35 switched on and turn up all the dials

1:37 that are Switched Off moving to our next pattern

1:40 a shifted j-shape our perceptron again outputs

1:43 a positive value instead of the desired negative value

1:47 so we again turn down all the switches that are on and turn up all the switches

1:51 that are off arriving at our fourth

1:53 and final example the current configuration of dials outputs

1:57 the correct positive value that we expect

1:59 for a in this case our procedure tells us

2:02 to leave our dials alone cycling back through

2:05 our patterns we see that our machine has

2:07 learned to correctly classify all four examples this procedure

2:11 was discovered in 1957 by the psychologist Frank

2:14 Rosen blot and is known as the perceptron

2:16 learning rule Rosen blot unveiled the approach

2:19 to the public in a press conference on July

2:21 7th 1958 the next day the New York Times

2:25 reported that the machine was expected to be

2:26 able to walk talk see write reproduce itself

2:30 and be conscious of its own existence Rosen

2:33 bl's perceptron is in some ways more sophisticated than

2:36 our machine its input grid was 20x 20

2:39 instead of 4x4 it had multiple artificial neurons

2:42 instead of just one and it used motors

2:44 to turn the dials so learning was entirely automatic

2:48 but the learning algorithm and operating principles are

2:50 the same rosenblatt's claims are grandiose but are

2:54 slowly coming true and before his untimely death

2:58 in 1971 Rosen blot was even working on multi-layer

3:01 architectures that closely resemble modern neural networks

3:05 but there was a problem with Rosen bl's design

3:07 and in fact all neural networks from this era

3:10 that nearly completely halted our modern neural

3:13 network driven approach to AI building the perceptron

3:17 machine for this video took quite a few

3:18 early morning design and soldering sessions on projects

3:22 like these I really like to have my morning

3:23 routine dialed in and this video sponsor ag1

3:27 is a key part of my routine a couple of years ago I found myself feeling extra

3:31 rundown and getting sick more often than usual

3:35 so I decided to have some detailed blood testing done to see if anything was off

3:39 my doctor found that my vitamin D levels were

3:41 very low and after taking supplements for a couple

3:44 of months I felt a huge difference this experience led me to have a broader look

3:49 at how I could optimize my nutrition and ag1 has been a terrific tool for me

3:54 I was able to replace a few separate

3:55 supplements with a single serving of ag1 each

3:58 morning the ingredient list is really impressive the biggest

4:01 benefits I noticed when taking ag1 are improved

4:04 energy and digestion I stopped my morning routine

4:07 over the holidays leading me to forget to take

4:09 ag1 and by the end of the break I found myself with less energy even though

4:13 I was getting more sleep ag1 is research

4:16 backed they use these cool machines for invitro

4:19 studies that simulate the digestive tract allowing for very

4:23 controlled study of a1's impact on the gut

4:25 microbiome the ag1 team also conduct studies

4:28 with human participants in a recent study 97%

4:32 of participants reported feeling more energy after taking

4:34 ag1 for one month you can get $20 off

4:38 your first subscription to ag1 by visiting drink

4:40 a1.com Welch laabs or by clicking the link

4:44 in the description below big thank you to ag1

4:47 for sponsoring this video now back to the perceptron

4:51 we've seen that our perceptron machine can quickly

4:53 learn to tell apart certain patterns but what

4:56 exactly can the perceptron do and not do

4:59 in 1962 Albert novakov proved mathematically that if

5:03 a configuration of dials exists that cleanly separates

5:06 a given set of examples the perceptron learning

5:08 rule is guaranteed to find it but do

5:11 cleanly separating configurations of dials exist for all types

5:14 of input patterns to get to the bottom

5:17 of this let's build an even simpler version

5:19 of the perceptron with just two inputs we

5:22 now have only four possible input patterns both switches

5:26 off one or the other switch on or both switches on note that although we only

5:31 have two inputs we have three dials the extra

5:34 dial is called bias and is not connected

5:36 to any of our switches but is effectively

5:38 always switched on the bias dial allows us

5:41 to directly add or subtract from the final value

5:43 that goes to our meter regardless of the current

5:46 switch configuration this is also why our full

5:48 machine has 17 dials instead of 16 now let's see that we want our perceptor

5:53 and to Output a positive value when either one

5:56 or both switches are on and a negative

5:58 value when both switches are off following our perceptron

6:01 learning rule our machine is able to successfully

6:04 learn these patterns in just three steps but what

6:09 about other assignments for our examples what if

6:12 we want the output to be positive when

6:13 either one of our switches is on and negative

6:16 when both switches are on or when both

6:18 switches are off following our same perceptron learning

6:21 rule we now get stuck in a loop where the machine never settles down on a viable

6:26 solution why is the perceptron able to learn

6:29 the first group of patterns but not the second we can make sense of what's going

6:33 on visually by plotting each input configuration

6:36 on a 2d grid where the x- axis represents

6:39 our first input value and our y- axis represents our second input value a switch

6:43 in the on position creates an input of plus

6:45 one volt so the configuration with both switches on is

6:49 represented by a point at 1 one on our grid an off switch represents an input

6:53 of minus1 volt so the configuration with both

6:56 switches off would show up at minus one minus

6:58 one on our grid and configurations with one switch on show up at 1 minus one

7:02 and minus one one now how does the mathematics

7:05 of our perceptron show up on our 2D grid each dial in our machine is effectively

7:10 multiplying each input value by a configurable weight

7:14 so our machine's output is equal to the weight value of our first dial time X

7:18 plus the weight value of our second dial time y plus the output of our bias

7:22 dial which doesn't depend on X or Y

7:25 we'll call this value B when our output value

7:28 is greater than zero our perceptron will classify

7:31 our input XY as positive now for what regions of our grid will the output

7:36 of our perceptron be greater than zero setting our output

7:40 greater to zero in our equation and solving

7:42 for y we get an equation for a straight

7:45 line with a slope of minus W1 over W2 and a y intercept of minus B

7:50 over W2 so our perceptron will classify all points on our grid above this line

7:55 as positive we can see this behavior in action

7:58 on the first example we trained our perceptron

8:00 on where we classified examples with either one

8:03 or both of our switches on as positive on our grid these are the points- one one

8:08 1 one and 1 minus one as the perceptron

8:11 learns and we update our weights our line

8:14 moves and quickly lands in a configuration where

8:16 all positive examples are on the positive side

8:18 of our decision boundary now what about the example

8:22 our perceptron failed to learn where we want

8:24 the machine to classify configurations as positive where

8:27 one but not both switches are

8:29 on on our grid these configurations correspond to the points

8:32 minus one one and 1 minus one watching

8:35 our perceptron try to learn this pattern we

8:37 can see our line jump around without being

8:39 able to settle in a location that cleanly

8:41 separates our examples this is because there's actually

8:45 no way to use a single line to separate

8:47 this pattern we'll always miss at least one

8:50 example said differently our data is not linearly

8:54 separable this example is known as the exclusive

8:56 or problem because the mapping from inputs to outputs

9:00 follows The Logical exclusive or function and is

9:03 one of the simplest examples of a nonlinearly

9:05 separable pattern this inability to learn a simple

9:08 exclusive or function was a major criticism of Rosen

9:11 blots perceptron and other early neural networks

9:15 the example may feel a bit contrived but linear

9:18 separability is a significant concern in real data

9:21 in Rosen bl's 1958 press conference he showed

9:24 an example of the perceptron learning to tell

9:26 apart examples with markings on the left versus the right side side of an image

9:30 these patterns are linearly separable however later that year

9:34 in an interview with the New Yorker Rosen

9:36 block claimed that the perceptron could tell apart images

9:39 of cats and dogs he likely meant in principle but if we try this out in Python

9:44 and images of cat and dog faces

9:45 the perceptron learning rule is able to learn some

9:48 basic patterns but is unstable and unable to beat

9:51 around a 20% error rate meaning this cat

9:53 and dog data set is not linearly separable

9:56 now as Rosen blot himself pointed out there is

9:59 a solution to the linear separability problem

10:02 but it comes with a catch our perceptron machine

10:04 implements a single artificial neuron which is limited

10:08 to creating a single linear decision boundary to recognize

10:11 patterns if we expand to a network

10:13 of these artificial neurons we can combine multiple linear

10:16 decision boundaries to learn more complex patterns it

10:20 turns out that a very simple network with just

10:23 three neurons across two layers can solve

10:25 the exclusive or problem it works by creating two

10:28 decision boundaries in the the first layer

10:30 and then combining these boundaries in the second layer

10:32 into a band on our grid that captures minus one one and 1us one while rejecting

10:37 minus1 minus one and 1 one however while

10:40 it was well known in the 1960s that multi-layer

10:43 neural networks could solve nonlinearly separable problems like

10:46 this in principle no one could find a suitable

10:48 algorithm for learning the weights the perceptron learning

10:51 rule that works so well for a single

10:53 neuron does not generalize to multi-layer networks

10:57 and while it's easy to manually figure out weights

10:59 to solve toy problems like exclusive ore real

11:01 problems like telling apart cats and dogs with multi-layer

11:04 neural networks requires a learning algorithm that can

11:06 simultaneously update all neuron weights based on real

11:09 data neural networks had hit a dead end in the mid 1960s two leaders in the AI

11:17 field Marvin Minsky and Seymour papert began circulating

11:20 pre-prints of a book they called perceptrons despite

11:24 its title the book generally takes a negative

11:26 view rigorously showing mathematically what perceptrons could not do

11:31 the cover of the book shows two figures

11:33 that the perceptron cannot tell apart these shapes

11:36 look similar but have different connectivities the top

11:39 shape is made from a single purple line and the bottom shape is made from two

11:43 like the exclusive or problem Rosen blots perceptron

11:46 cannot separate these patterns although modern neural networks

11:50 can this apparent dead end contributed to the rise

11:54 of completely different symbolic approaches to AI

11:56 in the 1960s and70s and a dramatic decline in neural

12:00 network research while Rosen blots percepton received most

12:04 of the attention at the time there were

12:06 a number of other groups working on systems

12:08 of artificial neurons at Stanford Bernard wdr's group

12:12 came Incredibly Close to solving the problem of training

12:14 multi-layer neural networks and discovering our modern approach

12:18 to AI on a Friday afternoon in the fall of 1959 woodro had his first meeting

12:24 with a new graduate student Ted Hoff as widrow

12:27 explained his group had been working on great

12:29 rent based methods for learning the weights in artificial

12:32 neurons the idea was that instead of following

12:34 an ad hoc method like the perceptron learning

12:37 rule it was possible to measure and minimize

12:40 the machine's error mathematically in our first two

12:43 input example where we wanted our percep Tron

12:45 to learn to classify inputs where either one

12:47 or both switches are on as positive we can

12:50 measure our systems error by comparing the machines

12:53 outputs to our Target outputs making a small

12:56 notation change we can write the output y

12:59 of our two input perceptron is w0* our first

13:02 input x0 plus W1* our second input X1+

13:06 B taking our first input pattern with both

13:09 switches off we want our perceptron to Output a value of y= minus1 if our dials

13:15 are set to min-1 1 and 1 for example this makes our output y hat equal

13:20 to -1* -1 plus -1* POS 1+ 1 which equals POS 1 however our Target value is

13:28 y= -1 so the difference or error between our Target and our output is y- y

13:34 hat= -2 from here wdr's group would Square

13:38 this error value this ensured the value was always

13:41 positive and made it easier to minimize so

13:44 our squared error is four now how does

13:47 our squared error change as we turn our dials

13:50 if we move our first dial from minus1 to minus 0.9 our overall output is now

13:55 equal to positive 0.9 making our error 1.9

13:59 and our squared error 3.61 moving our w0

14:03 dial step by step and plotting the results we see that our error value makes

14:07 a smooth parabolic shape reaching a minimum around w0

14:11 equals 1 of course we want to find

14:13 the best overall configuration of dials not just w0

14:17 if we vary w0 and W1 at the same time across a grid of values we end

14:22 up with a nice Bowl shape where the best

14:25 configuration of dials is located at the bottom

14:27 of the bowl now as we move to machines with more dials it becomes intractable

14:32 to test all configurations like this trying 10

14:35 values for each dial in our 16 input perceptron

14:38 would require an enormous 10 to the 17th

14:42 computations what's interesting here though is that we

14:44 don't actually need to know what our entire

14:46 eror landscape looks like to find the best solution

14:50 given a starting configuration of dials if we

14:52 just know which way is downhill on our error

14:54 landscape also known as the gradient we can

14:57 take incremental steps in this direction to reach the bottom of our bow up until

15:02 this point wdr's group estimated the gradient numerically given

15:06 our starting point of w0= minus1 W1= 1 and Bal 1 they would compute the error

15:12 at points in the neighborhood around each weight

15:16 for the w0 weight for example we can compute the error at w0= -1.1 and w0=

15:22 .9 to estimate the slope of the eror surface

15:26 as widra walked through this mathematics at his office

15:28 black with Hof they came across a powerful

15:31 new approach that is incredibly close to how

15:34 we train neural networks today but with one

15:36 critical exception instead of estimating the gradient

15:39 by Computing the error around each weight value What

15:42 If instead we use calculus to take the derivative

15:45 of the error function directly we can compute

15:48 the partial derivative of our error equation

15:50 with respect to our weight w0 by dropping down

15:53 the power of two and following the chain

15:55 rule giving -2* y- Y2 times the derivative

15:59 of y hat with respect to w0 writing out the full expression for y hat we can

16:05 see that only the w0 x0 term depends on w0 treating the input x0 as a constant

16:12 we can think of the system as a line with a slope of x0 so its

16:16 derivative is just x0 we're left with a simple

16:19 expression for de dw0- 2* y- y hat* x0 and we can compute similar expressions

16:27 for dw1 and Deb putting these pieces together we're

16:32 left with a simple equation that gives us

16:34 the full gradient telling us exactly which way is

16:37 downhill from a given starting point in our error

16:39 landscape as WID would later explain you

16:42 don't have to square anything or compute the actual

16:45 error the power of that compared to earlier

16:48 methods is just fantastic WID and Hoff had

16:51 found a method that appeared to be incredibly

16:53 efficient but would it actually work they had

16:57 to find out they quickly worked out a circuit

16:59 design but the university stock room was closed

17:02 by the end of their Friday afternoon meeting

17:04 and would not reopen until Monday the next

17:06 morning wdro and Hoff rushed to an electronic supply

17:09 store and over the weekend pieced together

17:11 an artificial neuron circuit very similar to ours

17:14 and by Sunday evening they were testing out

17:16 their new learning algorithm on various input patterns the new

17:20 algorithm worked incredibly well interestingly despite being

17:24 derived completely differently wdro and Hoff's algorithm updates

17:27 The Machine's weights and a very similar way

17:30 to the perceptron learning rule the only real difference

17:33 is that wdro and Hoff's algorithm which they

17:35 later called LMS includes multiplying the weight update

17:38 by the current error value so instead of turning

17:41 the knobs by the same fixed learning rate

17:43 at each step we turn them proportionally

17:46 to the error between the Target and current output

17:49 values the LMS algorithm can quickly solve the T&J

17:53 classification problem we solved earlier with the perceptron

17:55 learning rule it was later shown that although

17:58 the elements algorithm is not guaranteed to find

18:00 a solution if one exists it performs better in the all to Common case where

18:05 our examples are not linearly separable now as impressive

18:08 as the LMS algorithm is wdro and Hoff were not able to use it to address

18:13 the core issue that halted 1960s neural network

18:16 research as widra would later explain despite

18:20 their best efforts they were never able to adapt

18:22 LMS to train networks with multiple layers what's

18:26 Wild here is how close the algorithm they wrote

18:29 on the Blackboard that day is to the back

18:31 propagation algorithm that we use to train

18:33 modern systems like chat GPT returning to our simple

18:37 three neuron two-layer network from earlier that could

18:39 solve the nonlinearly separable exclusive or problem we

18:43 can write a set of equations that map

18:45 our inputs X to our outputs y hat just as woodro and Hoff did for a single

18:49 neuron following woodro and Hoff's approach we can

18:52 differentiate our error equation with respect to our weights

18:55 and follow the chain rule however when we

18:58 reach the boundary between our first and second

19:00 layers our calculus hits a brick wall the artificial

19:04 Neuron model that almost everyone used in the 50s

19:07 and 60s originally developed by Walter pittz

19:10 and Warren mullik in the 1940s assumes an All

19:13 or Nothing output if the sum of our weights

19:16 times our inputs is greater than zero

19:18 our artificial neuron fires and outputs a one

19:21 otherwise the neuron outputs a zero the slope

19:24 of this binary step activation function is zero

19:27 everywhere and undefined at the the origin so when

19:30 we reach this function in our chain rule

19:32 we get stuck the function effectively snaps our gradient

19:35 to zero visually the binary step activation

19:38 function turns our error landscape into flat plateaus

19:41 and infinitely steep Cliffs here's our error surface

19:45 is a function of the weights in our first

19:47 layer these flat Landscapes give gradient methods like

19:50 LMS no hope of incrementally working their way

19:53 downhill the solution to this problem in hindsight

19:56 is surprisingly simple all we have to to do

19:58 is replace the All or Nothing step activation

20:01 function with something less flat swapping our step

20:04 activation function for a sigmoid function our lost

20:07 landscape now looks like this we still have

20:10 somewhat flat regions but there's enough of a slope

20:13 in these regions to guide our solution

20:15 downhill here's the continuation of the LMS algorithm

20:18 extended through both layers solving the exclusive ore problem

20:23 in 1986 27 years after woodro and Hoff

20:26 discovered the LMS algorithm the David rumelhart Jeff

20:29 Hinton and Ronald Williams published this paper where

20:32 they present the modern back propagation algorithm that we

20:35 use today their derivation cites the LMS algorithm

20:38 which they call the Delta Rule and they proceed

20:41 with deriving back propagation as a generalization

20:43 of this rule using the chain Ru with sigmoid activation

20:46 functions to continue the gradient computation started by wro

20:49 and Hoff almost three decades before in 2020

20:54 open AI used back propagation to train gpt3

20:57 the largest neur Network ever created at the time

21:00 with a staggering 175 billion learnable weights

21:04 these weights are spread across 96 layers each made

21:07 of two compute blocks the second compute Block

21:10 in each layer is still called by the name

21:12 Frank Rosen block gave it almost 70 years

21:14 ago a multi-layer perceptron each multi-layer perceptron block

21:18 in gpt3 has two layers just like

21:21 our two-layer Network that solve the exclusive War problem

21:24 but with way more neurons around 50,000

21:27 in the first layer and 12,000 in the second

21:30 layer the first compute Block in each layer

21:32 implements a relatively recent idea called attention and arguably

21:36 still uses artificial neurons but in a more complex way one way to think about

21:41 the attention block is as a specialized multi-layer

21:43 perceptron where the weights are controlled by other perceptrons

21:47 allowing the data itself to control the weights

21:49 as it moves through the model each attention

21:52 Block in gpt3 effectively uses around 50,000 neurons

21:56 this makes for around 10 million artificial neurons across

21:59 gpt3 is 96 layers so we would need

22:03 10 million of our perceptron machines each with over

22:05 10,000 dials to implement gpt3 today gpt3 fits

22:10 on around 10 gpus GPT 4 is reportedly

22:14 around 10 times larger than gpt3 bringing our neuron

22:17 count to Something in the neighborhood of 100

22:20 million like our simple perceptron machine GPT 4

22:23 is in many ways a pattern recognizer its

22:26 enormous network of neurons allows it to learn

22:28 very complex patterns in language and use these patterns

22:32 to predict what text should come next it's

22:35 incredible that this Atomic unit the perceptron connected

22:38 in giant networks and trained with an extension

22:41 of the LMS algorithm would result in the most

22:44 intelligent systems we've been able to build so

22:46 far it's been almost 70 years now since

22:49 Frank rosenblau claimed that the perceptron would be

22:51 able to walk talk see write reproduce itself

22:55 and be conscious of its own existence he

22:58 clearly was missing some key details but time has

23:01 only proven Rosen blot more right about what

23:03 the humble perceptron can do we'll have to wait

23:06 and see if he was right about everything

23:12 big thanks to everyone who bought an imaginary

23:14 numbers book the first print run totally sold

23:17 out but the second print run just came

23:19 in and is available to order today the book

23:22 follows my imaginary number series including remon surfaces

23:26 and ends with new chapters on Oilers form

23:28 and Schrodinger's equation these chapters were a lot

23:31 of fun to put together I'm really happy

23:34 with how the figures and type setting came out

23:36 the book is printed on high quality heavyweight

23:38 paper with great detail and color reproduction whether

23:41 you're an expert looking for a different angle

23:43 on a topic you know well or just picking this stuff up for school work or fun

23:48 I really think you'll enjoy the book imaginary

23:51 numbers are such a deep and beautiful topic

23:53 get your book today at Welch labs.com resources come

Study with Looplines Download Captions Watch on YouTube