ChatGPT is made from 100 million of these [The Perceptron]
Welch Labs
0:00 this is a perceptron the machine shocked
0:02 the world in the 1950s by learning to recognize
0:05 patterns completely automatically and today the algorithm it
0:09 implements has become the core building block of AI
0:11 systems like chat GPT but why is
0:14 this the atomic unit of the intelligent systems we
0:16 have today the perceptron works by processing patterns
0:20 we input using these switches like this t-shape
0:23 switches in the up position output a positive
0:25 voltage and switches in the down position output
0:27 a negative voltage each switch is connected to an indicator
0:30 LED and then to one of these dials
0:33 rotating a dial multiplies the output of the switch
0:36 by the number shown on the dial and the meter shows the result of adding
0:40 the signals from all the dials together now
0:43 is there a way to configure our dials
0:45 such that the perceptron always outputs a positive
0:48 signal for t-shapes while outputting a negative value
0:51 for other types of shapes like this J shape note that our shapes won't be
0:55 in the same position each time remarkably it turns
0:58 out that if a configuration of that solves
1:00 our problem exists there is a simple procedure we
1:03 can follow that is guaranteed to find it
1:05 every time starting with this t-shape our meter
1:08 is showing a value close to zero but we want it to be positive in this case
1:13 our procedure tells us to turn all the knobs that are switch on to the right
1:17 by a constant value called The Learning rate
1:19 and to turn all the knobs that are Switched
1:21 Off to the left by the Learning rate our next pattern is a j-shape so we
1:26 want our machine to Output a negative value
1:28 however the current dial configuration outputs a positive
1:31 value in this case our procedure tells us to turn down all the dials that are
1:35 switched on and turn up all the dials
1:37 that are Switched Off moving to our next pattern
1:40 a shifted j-shape our perceptron again outputs
1:43 a positive value instead of the desired negative value
1:47 so we again turn down all the switches that are on and turn up all the switches
1:51 that are off arriving at our fourth
1:53 and final example the current configuration of dials outputs
1:57 the correct positive value that we expect
1:59 for a in this case our procedure tells us
2:02 to leave our dials alone cycling back through
2:05 our patterns we see that our machine has
2:07 learned to correctly classify all four examples this procedure
2:11 was discovered in 1957 by the psychologist Frank
2:14 Rosen blot and is known as the perceptron
2:16 learning rule Rosen blot unveiled the approach
2:19 to the public in a press conference on July
2:21 7th 1958 the next day the New York Times
2:25 reported that the machine was expected to be
2:26 able to walk talk see write reproduce itself
2:30 and be conscious of its own existence Rosen
2:33 bl's perceptron is in some ways more sophisticated than
2:36 our machine its input grid was 20x 20
2:39 instead of 4x4 it had multiple artificial neurons
2:42 instead of just one and it used motors
2:44 to turn the dials so learning was entirely automatic
2:48 but the learning algorithm and operating principles are
2:50 the same rosenblatt's claims are grandiose but are
2:54 slowly coming true and before his untimely death
2:58 in 1971 Rosen blot was even working on multi-layer
3:01 architectures that closely resemble modern neural networks
3:05 but there was a problem with Rosen bl's design
3:07 and in fact all neural networks from this era
3:10 that nearly completely halted our modern neural
3:13 network driven approach to AI building the perceptron
3:17 machine for this video took quite a few
3:18 early morning design and soldering sessions on projects
3:22 like these I really like to have my morning
3:23 routine dialed in and this video sponsor ag1
3:27 is a key part of my routine a couple of years ago I found myself feeling extra
3:31 rundown and getting sick more often than usual
3:35 so I decided to have some detailed blood testing done to see if anything was off
3:39 my doctor found that my vitamin D levels were
3:41 very low and after taking supplements for a couple
3:44 of months I felt a huge difference this experience led me to have a broader look
3:49 at how I could optimize my nutrition and ag1 has been a terrific tool for me
3:54 I was able to replace a few separate
3:55 supplements with a single serving of ag1 each
3:58 morning the ingredient list is really impressive the biggest
4:01 benefits I noticed when taking ag1 are improved
4:04 energy and digestion I stopped my morning routine
4:07 over the holidays leading me to forget to take
4:09 ag1 and by the end of the break I found myself with less energy even though
4:13 I was getting more sleep ag1 is research
4:16 backed they use these cool machines for invitro
4:19 studies that simulate the digestive tract allowing for very
4:23 controlled study of a1's impact on the gut
4:25 microbiome the ag1 team also conduct studies
4:28 with human participants in a recent study 97%
4:32 of participants reported feeling more energy after taking
4:34 ag1 for one month you can get $20 off
4:38 your first subscription to ag1 by visiting drink
4:40 a1.com Welch laabs or by clicking the link
4:44 in the description below big thank you to ag1
4:47 for sponsoring this video now back to the perceptron
4:51 we've seen that our perceptron machine can quickly
4:53 learn to tell apart certain patterns but what
4:56 exactly can the perceptron do and not do
4:59 in 1962 Albert novakov proved mathematically that if
5:03 a configuration of dials exists that cleanly separates
5:06 a given set of examples the perceptron learning
5:08 rule is guaranteed to find it but do
5:11 cleanly separating configurations of dials exist for all types
5:14 of input patterns to get to the bottom
5:17 of this let's build an even simpler version
5:19 of the perceptron with just two inputs we
5:22 now have only four possible input patterns both switches
5:26 off one or the other switch on or both switches on note that although we only
5:31 have two inputs we have three dials the extra
5:34 dial is called bias and is not connected
5:36 to any of our switches but is effectively
5:38 always switched on the bias dial allows us
5:41 to directly add or subtract from the final value
5:43 that goes to our meter regardless of the current
5:46 switch configuration this is also why our full
5:48 machine has 17 dials instead of 16 now let's see that we want our perceptor
5:53 and to Output a positive value when either one
5:56 or both switches are on and a negative
5:58 value when both switches are off following our perceptron
6:01 learning rule our machine is able to successfully
6:04 learn these patterns in just three steps but what
6:09 about other assignments for our examples what if
6:12 we want the output to be positive when
6:13 either one of our switches is on and negative
6:16 when both switches are on or when both
6:18 switches are off following our same perceptron learning
6:21 rule we now get stuck in a loop where the machine never settles down on a viable
6:26 solution why is the perceptron able to learn
6:29 the first group of patterns but not the second we can make sense of what's going
6:33 on visually by plotting each input configuration
6:36 on a 2d grid where the x- axis represents
6:39 our first input value and our y- axis represents our second input value a switch
6:43 in the on position creates an input of plus
6:45 one volt so the configuration with both switches on is
6:49 represented by a point at 1 one on our grid an off switch represents an input
6:53 of minus1 volt so the configuration with both
6:56 switches off would show up at minus one minus
6:58 one on our grid and configurations with one switch on show up at 1 minus one
7:02 and minus one one now how does the mathematics
7:05 of our perceptron show up on our 2D grid each dial in our machine is effectively
7:10 multiplying each input value by a configurable weight
7:14 so our machine's output is equal to the weight value of our first dial time X
7:18 plus the weight value of our second dial time y plus the output of our bias
7:22 dial which doesn't depend on X or Y
7:25 we'll call this value B when our output value
7:28 is greater than zero our perceptron will classify
7:31 our input XY as positive now for what regions of our grid will the output
7:36 of our perceptron be greater than zero setting our output
7:40 greater to zero in our equation and solving
7:42 for y we get an equation for a straight
7:45 line with a slope of minus W1 over W2 and a y intercept of minus B
7:50 over W2 so our perceptron will classify all points on our grid above this line
7:55 as positive we can see this behavior in action
7:58 on the first example we trained our perceptron
8:00 on where we classified examples with either one
8:03 or both of our switches on as positive on our grid these are the points- one one
8:08 1 one and 1 minus one as the perceptron
8:11 learns and we update our weights our line
8:14 moves and quickly lands in a configuration where
8:16 all positive examples are on the positive side
8:18 of our decision boundary now what about the example
8:22 our perceptron failed to learn where we want
8:24 the machine to classify configurations as positive where
8:27 one but not both switches are
8:29 on on our grid these configurations correspond to the points
8:32 minus one one and 1 minus one watching
8:35 our perceptron try to learn this pattern we
8:37 can see our line jump around without being
8:39 able to settle in a location that cleanly
8:41 separates our examples this is because there's actually
8:45 no way to use a single line to separate
8:47 this pattern we'll always miss at least one
8:50 example said differently our data is not linearly
8:54 separable this example is known as the exclusive
8:56 or problem because the mapping from inputs to outputs
9:00 follows The Logical exclusive or function and is
9:03 one of the simplest examples of a nonlinearly
9:05 separable pattern this inability to learn a simple
9:08 exclusive or function was a major criticism of Rosen
9:11 blots perceptron and other early neural networks
9:15 the example may feel a bit contrived but linear
9:18 separability is a significant concern in real data
9:21 in Rosen bl's 1958 press conference he showed
9:24 an example of the perceptron learning to tell
9:26 apart examples with markings on the left versus the right side side of an image
9:30 these patterns are linearly separable however later that year
9:34 in an interview with the New Yorker Rosen
9:36 block claimed that the perceptron could tell apart images
9:39 of cats and dogs he likely meant in principle but if we try this out in Python
9:44 and images of cat and dog faces
9:45 the perceptron learning rule is able to learn some
9:48 basic patterns but is unstable and unable to beat
9:51 around a 20% error rate meaning this cat
9:53 and dog data set is not linearly separable
9:56 now as Rosen blot himself pointed out there is
9:59 a solution to the linear separability problem
10:02 but it comes with a catch our perceptron machine
10:04 implements a single artificial neuron which is limited
10:08 to creating a single linear decision boundary to recognize
10:11 patterns if we expand to a network
10:13 of these artificial neurons we can combine multiple linear
10:16 decision boundaries to learn more complex patterns it
10:20 turns out that a very simple network with just
10:23 three neurons across two layers can solve
10:25 the exclusive or problem it works by creating two
10:28 decision boundaries in the the first layer
10:30 and then combining these boundaries in the second layer
10:32 into a band on our grid that captures minus one one and 1us one while rejecting
10:37 minus1 minus one and 1 one however while
10:40 it was well known in the 1960s that multi-layer
10:43 neural networks could solve nonlinearly separable problems like
10:46 this in principle no one could find a suitable
10:48 algorithm for learning the weights the perceptron learning
10:51 rule that works so well for a single
10:53 neuron does not generalize to multi-layer networks
10:57 and while it's easy to manually figure out weights
10:59 to solve toy problems like exclusive ore real
11:01 problems like telling apart cats and dogs with multi-layer
11:04 neural networks requires a learning algorithm that can
11:06 simultaneously update all neuron weights based on real
11:09 data neural networks had hit a dead end in the mid 1960s two leaders in the AI
11:17 field Marvin Minsky and Seymour papert began circulating
11:20 pre-prints of a book they called perceptrons despite
11:24 its title the book generally takes a negative
11:26 view rigorously showing mathematically what perceptrons could not do
11:31 the cover of the book shows two figures
11:33 that the perceptron cannot tell apart these shapes
11:36 look similar but have different connectivities the top
11:39 shape is made from a single purple line and the bottom shape is made from two
11:43 like the exclusive or problem Rosen blots perceptron
11:46 cannot separate these patterns although modern neural networks
11:50 can this apparent dead end contributed to the rise
11:54 of completely different symbolic approaches to AI
11:56 in the 1960s and70s and a dramatic decline in neural
12:00 network research while Rosen blots percepton received most
12:04 of the attention at the time there were
12:06 a number of other groups working on systems
12:08 of artificial neurons at Stanford Bernard wdr's group
12:12 came Incredibly Close to solving the problem of training
12:14 multi-layer neural networks and discovering our modern approach
12:18 to AI on a Friday afternoon in the fall of 1959 woodro had his first meeting
12:24 with a new graduate student Ted Hoff as widrow
12:27 explained his group had been working on great
12:29 rent based methods for learning the weights in artificial
12:32 neurons the idea was that instead of following
12:34 an ad hoc method like the perceptron learning
12:37 rule it was possible to measure and minimize
12:40 the machine's error mathematically in our first two
12:43 input example where we wanted our percep Tron
12:45 to learn to classify inputs where either one
12:47 or both switches are on as positive we can
12:50 measure our systems error by comparing the machines
12:53 outputs to our Target outputs making a small
12:56 notation change we can write the output y
12:59 of our two input perceptron is w0* our first
13:02 input x0 plus W1* our second input X1+
13:06 B taking our first input pattern with both
13:09 switches off we want our perceptron to Output a value of y= minus1 if our dials
13:15 are set to min-1 1 and 1 for example this makes our output y hat equal
13:20 to -1* -1 plus -1* POS 1+ 1 which equals POS 1 however our Target value is
13:28 y= -1 so the difference or error between our Target and our output is y- y
13:34 hat= -2 from here wdr's group would Square
13:38 this error value this ensured the value was always
13:41 positive and made it easier to minimize so
13:44 our squared error is four now how does
13:47 our squared error change as we turn our dials
13:50 if we move our first dial from minus1 to minus 0.9 our overall output is now
13:55 equal to positive 0.9 making our error 1.9
13:59 and our squared error 3.61 moving our w0
14:03 dial step by step and plotting the results we see that our error value makes
14:07 a smooth parabolic shape reaching a minimum around w0
14:11 equals 1 of course we want to find
14:13 the best overall configuration of dials not just w0
14:17 if we vary w0 and W1 at the same time across a grid of values we end
14:22 up with a nice Bowl shape where the best
14:25 configuration of dials is located at the bottom
14:27 of the bowl now as we move to machines with more dials it becomes intractable
14:32 to test all configurations like this trying 10
14:35 values for each dial in our 16 input perceptron
14:38 would require an enormous 10 to the 17th
14:42 computations what's interesting here though is that we
14:44 don't actually need to know what our entire
14:46 eror landscape looks like to find the best solution
14:50 given a starting configuration of dials if we
14:52 just know which way is downhill on our error
14:54 landscape also known as the gradient we can
14:57 take incremental steps in this direction to reach the bottom of our bow up until
15:02 this point wdr's group estimated the gradient numerically given
15:06 our starting point of w0= minus1 W1= 1 and Bal 1 they would compute the error
15:12 at points in the neighborhood around each weight
15:16 for the w0 weight for example we can compute the error at w0= -1.1 and w0=
15:22 .9 to estimate the slope of the eror surface
15:26 as widra walked through this mathematics at his office
15:28 black with Hof they came across a powerful
15:31 new approach that is incredibly close to how
15:34 we train neural networks today but with one
15:36 critical exception instead of estimating the gradient
15:39 by Computing the error around each weight value What
15:42 If instead we use calculus to take the derivative
15:45 of the error function directly we can compute
15:48 the partial derivative of our error equation
15:50 with respect to our weight w0 by dropping down
15:53 the power of two and following the chain
15:55 rule giving -2* y- Y2 times the derivative
15:59 of y hat with respect to w0 writing out the full expression for y hat we can
16:05 see that only the w0 x0 term depends on w0 treating the input x0 as a constant
16:12 we can think of the system as a line with a slope of x0 so its
16:16 derivative is just x0 we're left with a simple
16:19 expression for de dw0- 2* y- y hat* x0 and we can compute similar expressions
16:27 for dw1 and Deb putting these pieces together we're
16:32 left with a simple equation that gives us
16:34 the full gradient telling us exactly which way is
16:37 downhill from a given starting point in our error
16:39 landscape as WID would later explain you
16:42 don't have to square anything or compute the actual
16:45 error the power of that compared to earlier
16:48 methods is just fantastic WID and Hoff had
16:51 found a method that appeared to be incredibly
16:53 efficient but would it actually work they had
16:57 to find out they quickly worked out a circuit
16:59 design but the university stock room was closed
17:02 by the end of their Friday afternoon meeting
17:04 and would not reopen until Monday the next
17:06 morning wdro and Hoff rushed to an electronic supply
17:09 store and over the weekend pieced together
17:11 an artificial neuron circuit very similar to ours
17:14 and by Sunday evening they were testing out
17:16 their new learning algorithm on various input patterns the new
17:20 algorithm worked incredibly well interestingly despite being
17:24 derived completely differently wdro and Hoff's algorithm updates
17:27 The Machine's weights and a very similar way
17:30 to the perceptron learning rule the only real difference
17:33 is that wdro and Hoff's algorithm which they
17:35 later called LMS includes multiplying the weight update
17:38 by the current error value so instead of turning
17:41 the knobs by the same fixed learning rate
17:43 at each step we turn them proportionally
17:46 to the error between the Target and current output
17:49 values the LMS algorithm can quickly solve the T&J
17:53 classification problem we solved earlier with the perceptron
17:55 learning rule it was later shown that although
17:58 the elements algorithm is not guaranteed to find
18:00 a solution if one exists it performs better in the all to Common case where
18:05 our examples are not linearly separable now as impressive
18:08 as the LMS algorithm is wdro and Hoff were not able to use it to address
18:13 the core issue that halted 1960s neural network
18:16 research as widra would later explain despite
18:20 their best efforts they were never able to adapt
18:22 LMS to train networks with multiple layers what's
18:26 Wild here is how close the algorithm they wrote
18:29 on the Blackboard that day is to the back
18:31 propagation algorithm that we use to train
18:33 modern systems like chat GPT returning to our simple
18:37 three neuron two-layer network from earlier that could
18:39 solve the nonlinearly separable exclusive or problem we
18:43 can write a set of equations that map
18:45 our inputs X to our outputs y hat just as woodro and Hoff did for a single
18:49 neuron following woodro and Hoff's approach we can
18:52 differentiate our error equation with respect to our weights
18:55 and follow the chain rule however when we
18:58 reach the boundary between our first and second
19:00 layers our calculus hits a brick wall the artificial
19:04 Neuron model that almost everyone used in the 50s
19:07 and 60s originally developed by Walter pittz
19:10 and Warren mullik in the 1940s assumes an All
19:13 or Nothing output if the sum of our weights
19:16 times our inputs is greater than zero
19:18 our artificial neuron fires and outputs a one
19:21 otherwise the neuron outputs a zero the slope
19:24 of this binary step activation function is zero
19:27 everywhere and undefined at the the origin so when
19:30 we reach this function in our chain rule
19:32 we get stuck the function effectively snaps our gradient
19:35 to zero visually the binary step activation
19:38 function turns our error landscape into flat plateaus
19:41 and infinitely steep Cliffs here's our error surface
19:45 is a function of the weights in our first
19:47 layer these flat Landscapes give gradient methods like
19:50 LMS no hope of incrementally working their way
19:53 downhill the solution to this problem in hindsight
19:56 is surprisingly simple all we have to to do
19:58 is replace the All or Nothing step activation
20:01 function with something less flat swapping our step
20:04 activation function for a sigmoid function our lost
20:07 landscape now looks like this we still have
20:10 somewhat flat regions but there's enough of a slope
20:13 in these regions to guide our solution
20:15 downhill here's the continuation of the LMS algorithm
20:18 extended through both layers solving the exclusive ore problem
20:23 in 1986 27 years after woodro and Hoff
20:26 discovered the LMS algorithm the David rumelhart Jeff
20:29 Hinton and Ronald Williams published this paper where
20:32 they present the modern back propagation algorithm that we
20:35 use today their derivation cites the LMS algorithm
20:38 which they call the Delta Rule and they proceed
20:41 with deriving back propagation as a generalization
20:43 of this rule using the chain Ru with sigmoid activation
20:46 functions to continue the gradient computation started by wro
20:49 and Hoff almost three decades before in 2020
20:54 open AI used back propagation to train gpt3
20:57 the largest neur Network ever created at the time
21:00 with a staggering 175 billion learnable weights
21:04 these weights are spread across 96 layers each made
21:07 of two compute blocks the second compute Block
21:10 in each layer is still called by the name
21:12 Frank Rosen block gave it almost 70 years
21:14 ago a multi-layer perceptron each multi-layer perceptron block
21:18 in gpt3 has two layers just like
21:21 our two-layer Network that solve the exclusive War problem
21:24 but with way more neurons around 50,000
21:27 in the first layer and 12,000 in the second
21:30 layer the first compute Block in each layer
21:32 implements a relatively recent idea called attention and arguably
21:36 still uses artificial neurons but in a more complex way one way to think about
21:41 the attention block is as a specialized multi-layer
21:43 perceptron where the weights are controlled by other perceptrons
21:47 allowing the data itself to control the weights
21:49 as it moves through the model each attention
21:52 Block in gpt3 effectively uses around 50,000 neurons
21:56 this makes for around 10 million artificial neurons across
21:59 gpt3 is 96 layers so we would need
22:03 10 million of our perceptron machines each with over
22:05 10,000 dials to implement gpt3 today gpt3 fits
22:10 on around 10 gpus GPT 4 is reportedly
22:14 around 10 times larger than gpt3 bringing our neuron
22:17 count to Something in the neighborhood of 100
22:20 million like our simple perceptron machine GPT 4
22:23 is in many ways a pattern recognizer its
22:26 enormous network of neurons allows it to learn
22:28 very complex patterns in language and use these patterns
22:32 to predict what text should come next it's
22:35 incredible that this Atomic unit the perceptron connected
22:38 in giant networks and trained with an extension
22:41 of the LMS algorithm would result in the most
22:44 intelligent systems we've been able to build so
22:46 far it's been almost 70 years now since
22:49 Frank rosenblau claimed that the perceptron would be
22:51 able to walk talk see write reproduce itself
22:55 and be conscious of its own existence he
22:58 clearly was missing some key details but time has
23:01 only proven Rosen blot more right about what
23:03 the humble perceptron can do we'll have to wait
23:06 and see if he was right about everything
23:12 big thanks to everyone who bought an imaginary
23:14 numbers book the first print run totally sold
23:17 out but the second print run just came
23:19 in and is available to order today the book
23:22 follows my imaginary number series including remon surfaces
23:26 and ends with new chapters on Oilers form
23:28 and Schrodinger's equation these chapters were a lot
23:31 of fun to put together I'm really happy
23:34 with how the figures and type setting came out
23:36 the book is printed on high quality heavyweight
23:38 paper with great detail and color reproduction whether
23:41 you're an expert looking for a different angle
23:43 on a topic you know well or just picking this stuff up for school work or fun
23:48 I really think you'll enjoy the book imaginary
23:51 numbers are such a deep and beautiful topic
23:53 get your book today at Welch labs.com resources come