ADALINE & MADALINE Learning on the Net Input, Not the Output
The Adaptive Linear Neuron and its multi-unit successor — the first networks trained by true gradient descent, years before backpropagation.
From Perceptron to ADALINE
In 1960, Bernard Widrow and Ted Hoff (Stanford) introduced the ADAptive LINear Element — the same year as, but independently of, later perceptron refinements — solving a subtle but crucial flaw in the original perceptron learning rule.
The Key Idea 1960
The Perceptron Rule computes its error using the thresholded output: δ = target − ŷ, where ŷ ∈ {0,1}. This throws away information — two very different net inputs (say, net=0.01 and net=50) both produce ŷ=1 and hence zero learning signal once correctly classified.
ADALINE's insight: compute the error using the raw net input (before thresholding) instead. This makes the error a smooth, continuous, differentiable function of the weights — enabling true gradient descent.
📌 The activation function used during training is the identity (linear) function on the net input. A threshold is applied only afterward, at prediction time, to produce a binary class label.
Why It Matters Historically LINEAGE
ADALINE's Least-Mean-Squares (LMS) learning rule is mathematically identical in spirit to the Delta Rule used inside every layer of modern backpropagation. MADALINE (Many ADALINEs), introduced shortly after, was arguably the first practical multi-layer neural network ever deployed — used for decades in real-world adaptive filters and echo cancellation in telephone lines.
The ADALINE Model
Structurally similar to the perceptron, but with a critical difference in what gets compared to the target during training.
Architecture DEFINITION
net = Σᵢ wᵢxᵢ + b (training compares net to target) ŷ = sign(net) (used only at test time)
DIAGRAM — ADALINE UNIT
📌 This train-on-linear / predict-with-threshold split is the single most important structural difference from the perceptron, and it's what makes LMS a true gradient-descent method.
The Least-Mean-Squares (LMS) / Delta Rule
Derived directly from gradient descent on a squared-error cost function — the same derivation that underlies backpropagation.
Cost Function OBJECTIVE
For one training example, define the instantaneous squared error between the target and the net input (not the thresholded output):
E = ½ (t − net)² where net = Σᵢ wᵢxᵢ + b
Summed over all N training examples, the total cost is E_total = Σₙ ½(tₙ − netₙ)² — a smooth, convex bowl-shaped function of the weights (since net is linear in w).
Gradient Descent Derivation CALCULUS
// Differentiate E with respect to a single weight w_i
∂E/∂w_i = ∂E/∂net · ∂net/∂w_i
= −(t − net) · x_i
= −e · x_i // where e = (t − net)// Gradient descent: move w opposite the gradient
Δw_i = −η · ∂E/∂w_i = η · e · x_i
// Bias updates the same way, with x_0 = 1
Δb = η · e
Perceptron Rule vs. Delta Rule THE CRUCIAL DIFFERENCE
Property
Perceptron Rule
Delta / LMS Rule
Error uses
Thresholded output ŷ ∈ {0,1}
Raw net input (continuous)
Updates when
Only on misclassification
Every example (error rarely exactly 0)
Cost surface
Not well-defined / discontinuous
Smooth convex quadratic bowl
Guarantee
Converges only if linearly separable
Converges to minimum-MSE solution regardless
📌 Even on non-separable data, LMS still converges — to the weights that minimize total squared error, even if that solution misclassifies some points. This graceful degradation is a major practical advantage.
Interactive Lab — Train an ADALINE with the LMS Rule
Step through delta-rule updates on the AND truth table (using bipolar ±1 targets, the classical ADALINE convention) and watch the mean-squared error fall smoothly toward zero — not jump discretely like the perceptron trainer.
Bipolar AND Truth Table TARGET (±1)
x₁
x₂
t
−1
−1
−1
−1
1
−1
1
−1
−1
1
1
1
📌 ADALINE traditionally uses bipolar (±1) rather than binary (0/1) coding — it keeps the net input symmetric around zero and tends to train faster.
Trainer LIVE
CURRENT STATE
MEAN SQUARED ERROR OVER EPOCHS
TRAINING LOG
MADALINE — Many ADALINEs
A layer of ADALINE units feeding a fixed or trainable combination rule — the first practical multi-layer network, capable of solving non-linearly-separable problems like XOR.
Architecture STRUCTURE
MADALINE stacks several ADALINE units in a hidden layer, each producing a bipolar output (via sign), which then feed a combination/output unit — classically a fixed logic function (AND / OR / MAJORITY vote).
Hidden ADALINE units are each trained independently with the LMS rule against pre-assigned targets; the output combination logic (AND/OR/MAJORITY) has fixed, untrained weights. This sidesteps the "credit assignment problem" — nobody yet knew how to propagate error backward through a hidden layer.
MADALINE II / MRII (1980s) TRAINABLE HIDDEN LAYER
The MRII (Madaline Rule II) algorithm trains hidden-layer weights too, using a "minimal disturbance" heuristic: for each pattern, find the hidden unit whose net input is closest to zero (i.e., least confident), flip its output, and see if that reduces the overall error — adjust that unit's weights if so. A precursor to true backprop-trained hidden layers.
Solving XOR with MADALINE CLASSIC RESULT
Two hidden ADALINE units can each learn a different linear boundary; combined via majority vote / OR of one AND / etc., MADALINE recovers the same non-linear XOR decision region that a single-layer ADALINE (or perceptron) provably cannot reach.
x₁
x₂
z₁ (Adaline 1)
z₂ (Adaline 2)
ŷ = XOR
−1
−1
−1
−1
−1
−1
1
1
−1
1
1
−1
1
−1
1
1
1
−1
1
−1
Perceptron vs. ADALINE vs. MADALINE
Side-by-Side Comparison SUMMARY TABLE
Property
Perceptron
ADALINE
MADALINE
Layers
Single
Single
Multiple (hidden + output)
Learning rule
Perceptron rule (on ŷ)
LMS / Delta rule (on net)
LMS per unit (I) or MRII (II)
Error signal
Discrete
Continuous
Continuous per unit
Can solve XOR?
✗ No
✗ No
✓ Yes
Historical role
First learning neuron (1958)
First gradient-descent neuron (1960)
First practical multi-layer net
Python Example — ADALINE with the LMS Rule
adaline.py RUNNABLE
import numpy as np
# bipolar AND
X = np.array([[-1,-1],[-1,1],[1,-1],[1,1]])
t = np.array([-1,-1,-1,1])
w = np.zeros(2)
b = 0.0
eta = 0.1for epoch inrange(50):
mse = 0.0for xi, target inzip(X, t):
net = np.dot(w, xi) + b # linear, NOT thresholded
error = target - net
w += eta * error * xi # delta / LMS rule
b += eta * error
mse += error ** 2
mse /= len(X)
print(f"epoch {epoch}: w={w}, b={b:.3f}, mse={mse:.4f}")
if mse < 1e-4:
break# at test time, threshold with sign()defpredict(x):
return1if (np.dot(w, x) + b) >= 0else -1
📌 Notice: the loop computes error against net, not against a thresholded prediction. The threshold (sign) only appears in the separate predict() function used after training.