✦ ANN Module 02 · CI (CS-3030)

Perceptron Learning
From Fixed Weights to Learned Weights

Supervised training, the Perceptron Learning Rule, and hidden layers with backpropagation — plus a step-by-step interactive OR-classifier trainer.

Quick Recap — The Perceptron
A perceptron sums weighted inputs plus a bias, then applies a threshold activation to produce a binary output.
Perceptron Structure RECAP
Σ = Σᵢ (wᵢ xᵢ) + bias  →  f(x) = 1 if Σ ≥ 0, else 0  →  ŷ

Unlike the classical MCP neuron (equal weights, hand-set threshold), a perceptron's weights are learned from labeled training data via an iterative update rule.

DIAGRAM — PERCEPTRON STRUCTURE
x₁ x₂ ⋮ xₘ w₁ w₂ wₘ Σ+b step f(x) ŷ inputs summation + bias activation output
Single-Layer vs. Hidden-Layer Perceptron ARCHITECTURE

Single-Layer: input layer connects directly to a single output neuron — can only realize linearly separable functions (AND, OR, NOT).

Multi-Layer (with Hidden Layer): inputs connect to a hidden layer of neurons, which connect to the output — can realize non-linearly-separable functions like XOR.

DIAGRAM — SINGLE-LAYER vs. HIDDEN-LAYER
SINGLE-LAYER out WITH HIDDEN LAYER out input → output (linear only) input → hidden → output (XOR-capable)
The Single-Input Perceptron — Base Case
Before generalizing to m inputs, it's worth working through the simplest possible perceptron rigorously: one input, one weight, one bias.
Definition m = 1

For a single scalar input x with weight w and bias b, the perceptron computes a net input and applies the hard threshold (Heaviside) activation:

net = w·x + b     ŷ = f(net) = 1 if net ≥ 0, else 0
DIAGRAM — SINGLE-INPUT, SINGLE-OUTPUT NEURON
x w b net f(net) ŷ
📌 Example (from lecture notes): weight w=5, bias b=−3, input x=1 → net = 5×1−3 = 2. With bipolar sigmoid (s=0.5): f(2) = 1/(1+e^(−0.5×2)) ≈ 0.731.
Geometric Interpretation 1-D DECISION BOUNDARY

In m-dimensional input space, w·x + b = 0 defines an (m−1)-dimensional hyperplane separating the two classes. For m=1, that hyperplane collapses to a single point on the number line: x = −b/w. Everything to one side of that point classifies as 1, the other side as 0.

DECISION POINT ON THE NUMBER LINE (w=5, b=−3 → boundary at x=0.6)
−2 2 x=0.6 (boundary) ŷ=0 ŷ=1
Why It Matters FOUNDATION

The single-input case is the atomic unit of every larger network: an m-input perceptron is just m single-input perceptrons whose scalar products are summed before one shared activation. Understanding the 1-D decision point makes the m-dimensional hyperplane (Learning Rule section) far more intuitive — it is the exact same object, just embedded in higher dimensions.

ANN Working Models & Backpropagation
Breakthrough: Multi-Layer Perceptron (1986) HISTORY

Geoffrey Hinton, David Rumelhart, and Ronald Williams published "Learning representations by back-propagating errors," introducing two key ideas:

1. Backpropagation — a procedure to repeatedly adjust weights so as to minimize the difference between actual output and desired output.
2. Hidden Layers — neuron nodes stacked between inputs and outputs, letting the network learn more complex features (such as XOR logic).
Why Sigmoid, Not Step, for Learning MOTIVATION

With a hard step activation, a tiny change in any weight w can make the output jump suddenly from 0 to 1. We want gradual changes in weights to produce gradual changes in output — this is exactly what the smooth sigmoid enables.

z = Σᵢ wᵢxᵢ + bias,   σ(z) = 1 / (1 + e^(−z))
The Perceptron Learning Rule
Supervised Training PSEUDOCODE
1. Generate a training pair: input x = [x1 x2 ... xn] target output y_target 2. Present x to network, generate output y 3. Compare y with y_target to compute error 4. Adjust weights w to reduce error 5. Repeat 2-4 until done
Perceptron Learning Rule FORMULA
1. Initialize weights at random 2. For each pair (x, y_target): - compute output y - error δ = (y_target − y) - update: w_new = w_old + η·δ·x where η = learning rate 3. Repeat until δ = 0 (convergence)
The Rule, Formally KEY EQUATION
δ = (y_target − y)
w_new = w_old + η · δ · x

η (the learning rate or step size) determines how smoothly the learning process proceeds — too large and training oscillates, too small and it converges slowly.

Types of Learning Algorithms in ANNs
The Perceptron Learning Rule is just one instance of a broader taxonomy — how a network is allowed to see (or not see) the "right answer" during training defines the learning paradigm.
1. Supervised Learning HAS LABELS

Every training example (x, y_target) includes a known correct output. The network compares its prediction to the target and adjusts weights to reduce the error.

error = y_target − y  →  adjust weights

Examples:

• Perceptron Rule — discrete error δ=(target−output), updates only on misclassification.
• Delta Rule / LMS (Widrow–Hoff) — uses the continuous net input (before activation) to compute error, enabling gradient descent even when the output is linear (see ADALINE, Module 03).
• Backpropagation — generalizes the delta rule to multi-layer networks by propagating the output error backward through hidden layers via the chain rule.

2. Unsupervised Learning NO LABELS

No target output is provided — the network discovers structure (clusters, associations, compressed representations) purely from the statistics of the input data.

Examples:

• Hebbian Learning (1949) — "neurons that fire together, wire together": Δw = η·x·y, strengthening co-active connections without any external teacher.
• Competitive Learning / Self-Organizing Maps (SOM) — neurons compete for the right to respond to an input; the winner's weights move toward the input.
• Autoencoders — learn to reconstruct their own input through a compressed bottleneck.

3. Reinforcement Learning DELAYED REWARD

No labeled examples at all — an agent takes actions in an environment and receives a scalar reward signal, often delayed. It learns a policy that maximizes cumulative reward through trial and error, rather than directly minimizing a labeled error.

📌 Unlike supervised learning (told the exact correct answer every step) or unsupervised learning (told nothing), reinforcement learning is told only "how well it did," and must discover which of its actions were responsible.
Where the Perceptron Rule Fits SUMMARY
AlgorithmParadigmError signal usedUpdate rule
Perceptron RuleSupervisedDiscrete: (target − thresholded output)w += η·δ·x
Delta Rule (LMS)SupervisedContinuous: (target − net input)w += η·(t−net)·x
BackpropagationSupervisedChain-ruled gradient of lossw −= η·∂E/∂w
HebbianUnsupervisedNone — correlation onlyw += η·x·y
Demerits & Limitations of the (Single-Layer) Perceptron
Minsky & Papert's 1969 book "Perceptrons" formally proved several structural limits — findings that stalled neural network research for over a decade until the multi-layer perceptron + backpropagation revival of the 1980s.
1. Linear Separability Constraint FUNDAMENTAL

A single-layer perceptron can only correctly classify data that is linearly separable — separable by one straight line (2D), plane (3D), or hyperplane (n-D). XOR is the canonical counter-example: no single line separates its two classes.

XOR(x₁,x₂) = 1 ⟺ exactly one of x₁,x₂ is 1 — the two "1" points sit diagonally opposite, unreachable by any single line.
2. Convergence Only If Separable ⚠️

The Perceptron Convergence Theorem guarantees the learning rule finds a separating hyperplane in finite steps — but only if one exists. On non-separable data, the weights will oscillate forever and never settle.

3. Sensitivity to Feature Scaling ⚠️

Because updates are proportional to the raw input x, features on very different numeric scales can make training slow or unstable — unscaled inputs can dominate the weight updates disproportionately.

4. Hard, Non-Differentiable Activation ⚠️

The step function has zero gradient almost everywhere, so it cannot be used inside a gradient-descent/backpropagation pipeline — this is precisely why hidden-layer networks moved to smooth sigmoid/tanh/ReLU activations.

5. No Probabilistic Output ⚠️

The output is a hard 0/1 with no notion of confidence — there is no way to say "70% sure this is class 1," which limits its usefulness compared to sigmoid/softmax outputs.

The Resolution ✓

Every one of these limitations is addressed by moving to a multi-layer perceptron (hidden layers recover non-linear separability), a smooth activation function (enables gradient-based learning via backpropagation), and normalized inputs (stabilizes training). See the animated data-flow visualization below.

Multi-Layer Perceptron — Animated Forward Pass
Watch a real forward pass propagate step by step through an input layer, one hidden layer, and an output layer — with the exact arithmetic shown at each stage.
Network: 2 Inputs → 3 Hidden (sigmoid) → 1 Output (sigmoid) LIVE ANIMATION
NETWORK DIAGRAM — DATA FLOWS LEFT TO RIGHT
STEP-BY-STEP MATH (updates live as the animation plays)
Click "Animate Forward Pass" to begin.
Reading the Animation WHAT'S HAPPENING

1. Input layer lights up first — x₁ and x₂ enter unchanged.
2. Each hidden neuron computes its own weighted sum of both inputs plus its own bias, then squashes it through the sigmoid — nodes glow green as their value is computed.
3. The output neuron waits for all three hidden activations, computes its own weighted sum + bias, applies sigmoid, and produces the final ŷ.
4. This left-to-right sweep is exactly the forward pass — the same computation that, run in reverse with error gradients, becomes backpropagation.

Interactive Lab — Train an OR Perceptron
Step through the Perceptron Learning Rule epoch by epoch on the OR truth table, and watch the weights converge to a linear decision boundary.
OR Truth Table & Decision Boundary TARGET
xyout
000
011
101
111
Trainer LIVE
CURRENT STATE
DECISION BOUNDARY (w₁x + w₂y + b = 0)
TRAINING LOG
Python 3.6 Example — OR Perceptron
A minimal implementation of the perceptron learning rule matching the trainer lab above.
perceptron_or.py RUNNABLE
import numpy as np # OR truth table X = np.array([[0,0],[0,1],[1,0],[1,1]]) y = np.array([0,1,1,1]) w = np.zeros(2) b = 0.0 eta = 0.1 def step(z): return 1 if z >= 0 else 0 for epoch in range(10): errors = 0 for xi, target in zip(X, y): z = np.dot(w, xi) + b out = step(z) delta = target - out w += eta * delta * xi b += eta * delta errors += int(delta != 0) print(f"epoch {epoch}: w={w}, b={b}, errors={errors}") if errors == 0: break
📌 Run this and you'll see the perceptron converge in just a few epochs — OR is linearly separable, so convergence is guaranteed by the Perceptron Convergence Theorem.
Why This Wouldn't Work for XOR LIMITATION

XOR is not linearly separable — no single straight line (in 2D) can divide its four points into the correct classes. A single-layer perceptron will never converge on XOR; a hidden layer (Multi-Layer Perceptron) is required, trained via backpropagation.