- Calculate a weighted sum with a bias.
- Apply a ReLU activation.
- Trace information through network layers.
- Explain why nonlinearities matter.
- Distinguish forward computation and training updates.
Before you begin
Know linear predictions, loss, and evaluation. Basic arithmetic is sufficient; calculus is not required here.
A neuron is a calculation
Imagine deciding whether sensor readings deserve attention. A small artificial neuron can combine several numerical inputs into one output. Multiply each input by a weight, add the results and a bias, then apply an activation function. This calculation is the starting point for many neural networks.
The word neuron is an analogy inspired by biology. The mechanism in this lesson is not a faithful simulation of a human brain cell. It is an adjustable mathematical function. Keeping the arithmetic visible prevents the analogy from replacing understanding.
Weights control how inputs contribute. A positive weight can increase the sum as its input grows; a negative weight can decrease it. Bias shifts the result independently of the current input values.
Calculate one complete example
Use inputs x1 = 2 and x2 = 3, weights w1 = 0.5 and w2 = −1, and bias b = 2. The weighted sum is 2 × 0.5 + 3 × (−1) + 2 = 0. Apply ReLU, a function returning the larger of zero and its input, and the output remains zero.
Now change x1 to 4 while keeping everything else fixed. The sum becomes 4 × 0.5 − 3 + 2 = 1, and ReLU outputs 1. You have changed an input, not trained the neuron. Training would adjust weights or bias using data and a loss.
ReLU is one activation choice. Sigmoid maps a value into a range between zero and one; other functions serve other purposes. An output between zero and one is not automatically a well-calibrated probability. Its interpretation depends on model design and evaluation.
x1 × w1 ─┐
├→ sum + bias → activation → output
x2 × w2 ─┘Connect neurons into layers
A layer applies several such computations to incoming values. Its outputs become inputs to the next layer. An input layer represents supplied features, hidden layers perform intermediate transformations, and an output layer produces values appropriate to the task.
A regression network may output one number. A multiclass classifier may output one score per category, followed by a transformation such as softmax and a decision rule. Neither the number of layers nor the shape of a diagram alone tells you whether the network solves the intended problem.
Intermediate values are learned representations. They can be useful without corresponding to concepts that people can name easily. Avoid claiming that a particular layer always detects wheels or understands danger unless you have evidence for that specific model.
Why nonlinear functions matter
Stacking only linear transformations still produces an overall linear transformation, with biases giving an affine form. More layers alone would not create the desired range of nonlinear relationships. Activation functions break that restriction.
Consider points inside a circular boundary versus points outside it. A single straight dividing line cannot separate the two groups perfectly. A network with appropriate nonlinear transformations can represent a more flexible boundary. Capacity does not guarantee learning that boundary from inadequate data.
Google's neural-network module introduces this motivation. The practical lesson is that architecture determines what relationships can be represented, while training and data determine which behavior is actually obtained.
Training connects the output error to parameters
A forward pass computes outputs from inputs. A loss compares those outputs with the training objective. Backpropagation efficiently calculates how that loss changes with parameters through the chain of operations. An optimizer uses those gradients to update values.
Backpropagation calculates gradients; it is not itself a guarantee that the next model will be best. Learning rate, data scale, architecture, and optimization details matter. Networks can underfit, overfit, or learn shortcuts like other models.
For a beginner project, compare a network with a simpler baseline and inspect errors. Choose a network because evidence or task structure supports it, not because its name sounds more intelligent. Small arithmetic exercises are a sound beginning before you use a larger library.
Important terms
- Neuron
- A weighted-sum computation followed by an activation.
- Weight
- A learned coefficient multiplying an input.
- Bias
- A learned additive offset.
- Activation
- A function applied to a neuron's pre-activation value.
- Hidden layer
- An intermediate set of transformations.
- Backpropagation
- An efficient way to calculate gradients through a network.
Mini project: Build a two-neuron table
- Use the worked neuron's weights and bias for x1 values 0, 2, and 4 with x2 fixed at 3.
- Calculate sums −1, 0, and 1, then ReLU outputs 0, 0, and 1.
- Create a second neuron with weights 1 and 0 and bias −1.
- Calculate its outputs on the same inputs.
- Finish with two output columns and explain what information differs between them.
Common mistakes and debugging
- Confusing changing inputs with training: training changes parameters.
- Treating neurons as biological replicas: explain the actual calculation.
- Adding layers without nonlinearities and expecting arbitrary curved behavior: inspect the composed functions.
- Assuming more capacity guarantees generalization: retain suitable validation.
Independent challenge
Change the first neuron's bias to zero. Recalculate the table and explain how the activation threshold moves.
Check your understanding: 10 questions
What is the first example's weighted sum?
What does ReLU do to −2?
What changes during training?
Why use nonlinear activations?
What does backpropagation calculate?
In your own words, what does “Neuron” mean?
In your own words, what does “Weight” mean?
In your own words, what does “Bias” mean?
In your own words, what does “Activation” mean?
In your own words, what does “Hidden layer” mean?
Quiz answers
Reveal all 10 answers after your attempt
- Zero: 1 − 3 + 2 equals 0.
- It outputs zero.
- Parameters such as weights and biases are updated.
- They allow layered models to represent relationships beyond one overall linear transformation.
- Gradients relating the loss to parameters through the network.
- A weighted-sum computation followed by an activation.
- A learned coefficient multiplying an input.
- A learned additive offset.
- A function applied to a neuron's pre-activation value.
- An intermediate set of transformations.
Summary
Neural networks combine weighted sums and activations across layers. Their capabilities come from architecture, learned parameters, suitable data, and measured generalization.
Continue learning
ML07 applies learned representations to images and distinguishes classification, detection, and tracking.
- AI vs Machine Learning vs Deep Learning
- How Machine-Learning Models Learn From Error
- Computer Vision: How Machines Understand Images
- How to Design and Improve a Machine-Learning Project
Sources and further reading
Prepared 2026-09-18. Draft — arithmetic checked; network concepts source checked