Project: a neural network from scratch
Train a hidden layer to solve a pattern that one linear neuron cannot represent.
- Implement a two-layer forward pass
- Backpropagate through hidden units
- Verify a learned nonlinear boundary
A small nonlinear challenge
XOR is positive when exactly one of two binary inputs is one. The positive points lie on opposite corners of a square, so a single straight boundary cannot separate them from the negative points. A hidden layer with nonlinear activations can represent the pattern.
This project uses four tanh hidden units and one sigmoid output. Binary cross-entropy gives an output-logit gradient of prediction minus target. The hidden derivative multiplies the upstream gradient by one minus the squared tanh output.
Check the mechanics and the claim
Accumulate all gradients for a batch before updating weights. Each backward pass must use the same weights as its forward pass. A seeded initialization makes this educational experiment repeatable.
The four possible XOR inputs are all used for training and verification. This checks learning the truth table; it is not a held-out generalization experiment. Do not report its success as evidence of broader reasoning ability. To study generalization, define a larger task with separate examples.
A small experiment you can run.
The final assertions verify all four truth-table outputs. The hidden layer gives the network a nonlinear decision surface; the training loop optimizes it.
import math, random
rng = random.Random(17)
data = [([0., 0.], 0.), ([0., 1.], 1.), ([1., 0.], 1.), ([1., 1.], 0.)]
w1 = [[rng.uniform(-1, 1) for _ in range(2)] for _ in range(4)]
b1 = [0.] * 4
w2 = [rng.uniform(-1, 1) for _ in range(4)]
b2 = 0.
def forward(x):
h = [math.tanh(sum(w*t for w, t in zip(row, x))+b) for row, b in zip(w1, b1)]
p = 1/(1+math.exp(-(sum(w*v for w, v in zip(w2, h))+b2)))
return h, p
for epoch in range(5000):
gw1, gb1, gw2, gb2 = [[0., 0.] for _ in range(4)], [0.]*4, [0.]*4, 0.
for x, y in data:
h, p = forward(x)
dz = p-y
gb2 += dz
for j in range(4):
gw2[j] += dz*h[j]
dh = dz*w2[j]*(1-h[j]*h[j])
gb1[j] += dh
for k in range(2):
gw1[j][k] += dh*x[k]
rate = 0.3/len(data)
b2 -= rate*gb2
for j in range(4):
w2[j] -= rate*gw2[j]
b1[j] -= rate*gb1[j]
for k in range(2):
w1[j][k] -= rate*gw1[j][k]
for x, y in data:
p = forward(x)[1]
print(x, "target", int(y), "score", round(p, 3))
assert int(p >= 0.5) == int(y)
Save the file, open your terminal in that folder, and run python project-neural-network-xor.py. Use python3 or py if required by your installation. Setup guide
With the supplied seed and settings, all four XOR predictions pass the assertions.
Inspect the hidden representation.
- Print the four hidden activations for every input after training.
- Reduce the hidden layer to one unit and adapt all loop sizes.
- Compare convergence and describe what the hidden layer contributes.
Compare with a suggested solution
Hidden units produce different nonlinear responses to the input corners. Changing width changes representational capacity and optimization. Keep the same evaluation claim: learning four known cases is not evidence about an unrelated task.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
Python SQLite modulePython HTTP server: limitations