Unlike the classical MCP neuron (equal weights, hand-set threshold), a perceptron's weights are learned from labeled training data via an iterative update rule.
Single-Layer: input layer connects directly to a single output neuron — can only realize linearly separable functions (AND, OR, NOT).
Multi-Layer (with Hidden Layer): inputs connect to a hidden layer of neurons, which connect to the output — can realize non-linearly-separable functions like XOR.
For a single scalar input x with weight w and bias b, the perceptron computes a net input and applies the hard threshold (Heaviside) activation:
In m-dimensional input space, w·x + b = 0 defines an (m−1)-dimensional hyperplane separating the two classes. For m=1, that hyperplane collapses to a single point on the number line: x = −b/w. Everything to one side of that point classifies as 1, the other side as 0.
The single-input case is the atomic unit of every larger network: an m-input perceptron is just m single-input perceptrons whose scalar products are summed before one shared activation. Understanding the 1-D decision point makes the m-dimensional hyperplane (Learning Rule section) far more intuitive — it is the exact same object, just embedded in higher dimensions.
Geoffrey Hinton, David Rumelhart, and Ronald Williams published "Learning representations by back-propagating errors," introducing two key ideas:
2. Hidden Layers — neuron nodes stacked between inputs and outputs, letting the network learn more complex features (such as XOR logic).
With a hard step activation, a tiny change in any weight w can make the output jump suddenly from 0 to 1. We want gradual changes in weights to produce gradual changes in output — this is exactly what the smooth sigmoid enables.
η (the learning rate or step size) determines how smoothly the learning process proceeds — too large and training oscillates, too small and it converges slowly.
Every training example (x, y_target) includes a known correct output. The network compares its prediction to the target and adjusts weights to reduce the error.
Examples:
• Perceptron Rule — discrete error δ=(target−output), updates only on misclassification.
• Delta Rule / LMS (Widrow–Hoff) — uses the continuous net input (before activation) to compute error, enabling gradient descent even when the output is linear (see ADALINE, Module 03).
• Backpropagation — generalizes the delta rule to multi-layer networks by propagating the output error backward through hidden layers via the chain rule.
No target output is provided — the network discovers structure (clusters, associations, compressed representations) purely from the statistics of the input data.
Examples:
• Hebbian Learning (1949) — "neurons that fire together, wire together": Δw = η·x·y, strengthening co-active connections without any external teacher.
• Competitive Learning / Self-Organizing Maps (SOM) — neurons compete for the right to respond to an input; the winner's weights move toward the input.
• Autoencoders — learn to reconstruct their own input through a compressed bottleneck.
No labeled examples at all — an agent takes actions in an environment and receives a scalar reward signal, often delayed. It learns a policy that maximizes cumulative reward through trial and error, rather than directly minimizing a labeled error.
| Algorithm | Paradigm | Error signal used | Update rule |
| Perceptron Rule | Supervised | Discrete: (target − thresholded output) | w += η·δ·x |
| Delta Rule (LMS) | Supervised | Continuous: (target − net input) | w += η·(t−net)·x |
| Backpropagation | Supervised | Chain-ruled gradient of loss | w −= η·∂E/∂w |
| Hebbian | Unsupervised | None — correlation only | w += η·x·y |
A single-layer perceptron can only correctly classify data that is linearly separable — separable by one straight line (2D), plane (3D), or hyperplane (n-D). XOR is the canonical counter-example: no single line separates its two classes.
The Perceptron Convergence Theorem guarantees the learning rule finds a separating hyperplane in finite steps — but only if one exists. On non-separable data, the weights will oscillate forever and never settle.
Because updates are proportional to the raw input x, features on very different numeric scales can make training slow or unstable — unscaled inputs can dominate the weight updates disproportionately.
The step function has zero gradient almost everywhere, so it cannot be used inside a gradient-descent/backpropagation pipeline — this is precisely why hidden-layer networks moved to smooth sigmoid/tanh/ReLU activations.
The output is a hard 0/1 with no notion of confidence — there is no way to say "70% sure this is class 1," which limits its usefulness compared to sigmoid/softmax outputs.
Every one of these limitations is addressed by moving to a multi-layer perceptron (hidden layers recover non-linear separability), a smooth activation function (enables gradient-based learning via backpropagation), and normalized inputs (stabilizes training). See the animated data-flow visualization below.
1. Input layer lights up first — x₁ and x₂ enter unchanged.
2. Each hidden neuron computes its own weighted sum of both inputs plus its own bias, then squashes it through the sigmoid — nodes glow green as their value is computed.
3. The output neuron waits for all three hidden activations, computes its own weighted sum + bias, applies sigmoid, and produces the final ŷ.
4. This left-to-right sweep is exactly the forward pass — the same computation that, run in reverse with error gradients, becomes backpropagation.
| x | y | out |
| 0 | 0 | 0 |
| 0 | 1 | 1 |
| 1 | 0 | 1 |
| 1 | 1 | 1 |
XOR is not linearly separable — no single straight line (in 2D) can divide its four points into the correct classes. A single-layer perceptron will never converge on XOR; a hidden layer (Multi-Layer Perceptron) is required, trained via backpropagation.