✦ ANN Module 03 · Widrow & Hoff, 1960

ADALINE & MADALINE
Learning on the Net Input, Not the Output

The Adaptive Linear Neuron and its multi-unit successor — the first networks trained by true gradient descent, years before backpropagation.

From Perceptron to ADALINE
In 1960, Bernard Widrow and Ted Hoff (Stanford) introduced the ADAptive LINear Element — the same year as, but independently of, later perceptron refinements — solving a subtle but crucial flaw in the original perceptron learning rule.
The Key Idea 1960

The Perceptron Rule computes its error using the thresholded output: δ = target − ŷ, where ŷ ∈ {0,1}. This throws away information — two very different net inputs (say, net=0.01 and net=50) both produce ŷ=1 and hence zero learning signal once correctly classified.

ADALINE's insight: compute the error using the raw net input (before thresholding) instead. This makes the error a smooth, continuous, differentiable function of the weights — enabling true gradient descent.

📌 The activation function used during training is the identity (linear) function on the net input. A threshold is applied only afterward, at prediction time, to produce a binary class label.
Why It Matters Historically LINEAGE

ADALINE's Least-Mean-Squares (LMS) learning rule is mathematically identical in spirit to the Delta Rule used inside every layer of modern backpropagation. MADALINE (Many ADALINEs), introduced shortly after, was arguably the first practical multi-layer neural network ever deployed — used for decades in real-world adaptive filters and echo cancellation in telephone lines.

The ADALINE Model
Structurally similar to the perceptron, but with a critical difference in what gets compared to the target during training.
Architecture DEFINITION
net = Σᵢ wᵢxᵢ + b     (training compares net to target)     ŷ = sign(net)  (used only at test time)
DIAGRAM — ADALINE UNIT
x₁ x₂ ⋮ xₙ bias x₀=+1 w₁ w₂ wₙ Σ net identity (train) net compared to target t → error e = t − net → drives Δw (LMS rule) sign(·) (test) ŷ training uses the GREEN path · inference uses the AMBER path
📌 This train-on-linear / predict-with-threshold split is the single most important structural difference from the perceptron, and it's what makes LMS a true gradient-descent method.
The Least-Mean-Squares (LMS) / Delta Rule
Derived directly from gradient descent on a squared-error cost function — the same derivation that underlies backpropagation.
Cost Function OBJECTIVE

For one training example, define the instantaneous squared error between the target and the net input (not the thresholded output):

E = ½ (t − net)²     where   net = Σᵢ wᵢxᵢ + b

Summed over all N training examples, the total cost is E_total = Σₙ ½(tₙ − netₙ)² — a smooth, convex bowl-shaped function of the weights (since net is linear in w).

Gradient Descent Derivation CALCULUS
// Differentiate E with respect to a single weight w_i ∂E/∂w_i = ∂E/∂net · ∂net/∂w_i = −(t − net) · x_i = −e · x_i // where e = (t − net) // Gradient descent: move w opposite the gradient Δw_i = −η · ∂E/∂w_i = η · e · x_i // Bias updates the same way, with x_0 = 1 Δb = η · e
w_new = w_old + η·(t − net)·x     b_new = b_old + η·(t − net)
Perceptron Rule vs. Delta Rule THE CRUCIAL DIFFERENCE
PropertyPerceptron RuleDelta / LMS Rule
Error usesThresholded output ŷ ∈ {0,1}Raw net input (continuous)
Updates whenOnly on misclassificationEvery example (error rarely exactly 0)
Cost surfaceNot well-defined / discontinuousSmooth convex quadratic bowl
GuaranteeConverges only if linearly separableConverges to minimum-MSE solution regardless
📌 Even on non-separable data, LMS still converges — to the weights that minimize total squared error, even if that solution misclassifies some points. This graceful degradation is a major practical advantage.
Interactive Lab — Train an ADALINE with the LMS Rule
Step through delta-rule updates on the AND truth table (using bipolar ±1 targets, the classical ADALINE convention) and watch the mean-squared error fall smoothly toward zero — not jump discretely like the perceptron trainer.
Bipolar AND Truth Table TARGET (±1)
x₁x₂t
−1−1−1
−11−1
1−1−1
111
📌 ADALINE traditionally uses bipolar (±1) rather than binary (0/1) coding — it keeps the net input symmetric around zero and tends to train faster.
Trainer LIVE
CURRENT STATE
MEAN SQUARED ERROR OVER EPOCHS
TRAINING LOG
MADALINE — Many ADALINEs
A layer of ADALINE units feeding a fixed or trainable combination rule — the first practical multi-layer network, capable of solving non-linearly-separable problems like XOR.
Architecture STRUCTURE

MADALINE stacks several ADALINE units in a hidden layer, each producing a bipolar output (via sign), which then feed a combination/output unit — classically a fixed logic function (AND / OR / MAJORITY vote).

DIAGRAM — MADALINE (2 HIDDEN ADALINES + MAJORITY OUTPUT)
x₁ x₂ ADALINE z₁=sign(net₁) ADALINE z₂=sign(net₂) MAJORITY (or fixed AND/OR) ŷ hidden ADALINE layer fixed combination layer
MADALINE I (1960s) ORIGINAL

Hidden ADALINE units are each trained independently with the LMS rule against pre-assigned targets; the output combination logic (AND/OR/MAJORITY) has fixed, untrained weights. This sidesteps the "credit assignment problem" — nobody yet knew how to propagate error backward through a hidden layer.

MADALINE II / MRII (1980s) TRAINABLE HIDDEN LAYER

The MRII (Madaline Rule II) algorithm trains hidden-layer weights too, using a "minimal disturbance" heuristic: for each pattern, find the hidden unit whose net input is closest to zero (i.e., least confident), flip its output, and see if that reduces the overall error — adjust that unit's weights if so. A precursor to true backprop-trained hidden layers.

Solving XOR with MADALINE CLASSIC RESULT

Two hidden ADALINE units can each learn a different linear boundary; combined via majority vote / OR of one AND / etc., MADALINE recovers the same non-linear XOR decision region that a single-layer ADALINE (or perceptron) provably cannot reach.

x₁x₂z₁ (Adaline 1)z₂ (Adaline 2)ŷ = XOR
−1−1−1−1−1
−111−11
1−11−11
11−11−1
Perceptron vs. ADALINE vs. MADALINE
Side-by-Side Comparison SUMMARY TABLE
PropertyPerceptronADALINEMADALINE
LayersSingleSingleMultiple (hidden + output)
Learning rulePerceptron rule (on ŷ)LMS / Delta rule (on net)LMS per unit (I) or MRII (II)
Error signalDiscreteContinuousContinuous per unit
Can solve XOR?✗ No✗ No✓ Yes
Historical roleFirst learning neuron (1958)First gradient-descent neuron (1960)First practical multi-layer net
Python Example — ADALINE with the LMS Rule
adaline.py RUNNABLE
import numpy as np # bipolar AND X = np.array([[-1,-1],[-1,1],[1,-1],[1,1]]) t = np.array([-1,-1,-1,1]) w = np.zeros(2) b = 0.0 eta = 0.1 for epoch in range(50): mse = 0.0 for xi, target in zip(X, t): net = np.dot(w, xi) + b # linear, NOT thresholded error = target - net w += eta * error * xi # delta / LMS rule b += eta * error mse += error ** 2 mse /= len(X) print(f"epoch {epoch}: w={w}, b={b:.3f}, mse={mse:.4f}") if mse < 1e-4: break # at test time, threshold with sign() def predict(x): return 1 if (np.dot(w, x) + b) >= 0 else -1
📌 Notice: the loop computes error against net, not against a thresholded prediction. The threshold (sign) only appears in the separate predict() function used after training.