What you are looking at
Each dot is a training example with two numbers, its x and y position, and a label: indigo or amber. The neural network's job is to learn a rule that predicts the colour from the position. The shaded background shows the network's current prediction for every point on the plane: the stronger the colour, the more confident it is. When you press Train, the network repeatedly looks at the examples, measures how wrong it is and nudges its internal numbers, called weights, to be a little less wrong. Over a few seconds the background reshapes itself until it separates the dots.
How the network works
This is a multilayer perceptron written from scratch in plain JavaScript, with no libraries and no server. The input layer takes x and y. Each hidden neuron computes a weighted sum of the previous layer plus a bias and passes it through an activation function, such as tanh, which bends straight lines into curves. The output neuron uses a sigmoid to give a probability between 0 and 1. Training uses backpropagation: the error at the output is passed backwards through the layers using the chain rule, giving the gradient of the loss with respect to every weight. Mini-batch gradient descent then moves each weight a small step, set by the learning rate, in the direction that reduces the cross-entropy loss.
Experiments to try
- XOR with no hidden layer. Enter a single hidden layer of 1 neuron and see it fail: one neuron can only draw a straight line, and XOR needs two. Then try 2 or 4 neurons.
- The spiral. It needs more capacity. Try two layers of 12 to 16 neurons and be patient.
- Learning rate. Too low and progress is slow; too high and the loss jumps around or the boundary flickers.
- Noise and regularisation. Add noise so classes overlap, then compare training with and without L2 regularisation, which keeps weights small and boundaries smoother.
- Your own data. Clear the points and click to place your own, choosing the class first.
Why it matters
The networks behind image recognition and chatbots are vastly larger, but they learn with the same core recipe you can watch here: a differentiable model, a loss function and gradient descent. Seeing a boundary form, overfit noisy data or get stuck builds intuition that equations alone do not. Everything runs on your device, so it works offline once the page has loaded.