Try: Switch from Sigmoid to ReLU with 10+ layers. Watch the vanishing gradient problem disappear. Neuron size and color represent gradient magnitude.