Before delving into deep architectures, the linear unit serves as our foundational primitive. A set of linear units operating in parallel constitutes a fully connected (dense) layer; stacking multiple dense layers with non-linear activation functions enables the construction of deep neural networks (DNNs).
Figure 1:Simple perceptron with one input and one output.
The objective of simple linear regression is to predict the target data based on the input data :
where represents the unknown underlying function, and denotes irreducible noise independent of , conventionally distributed as .
Given that the true function is unknown to us, we can estimate by assuming that is approximately linear:
where is the predicted response.
Purpose of this Notebook:
Synthesize a 1D linear regression dataset.
Implement a single-unit Perceptron from first principles.
Derive analytical gradients via the multivariable chain rule.
Implement vectorized mini-batch gradient descent.
Train the custom model and monitor convergence.
Benchmark and numerically verify equivalence against PyTorch’s
nn.Linear.
Setup¶
print("Start package installation...")Start package installation...
%%capture
%pip install torch
%pip install scikit-learn
%pip install matplotlibprint("Packages installed successfully!")Packages installed successfully!
from platform import python_version
import torch
from torch import nn
python_version(), torch.__version__('3.14.7', '2.12.0+cpu')# Set seeds for reproducibility
import random
import numpy as np
SEED: int = 11
random.seed(SEED)
np.random.seed(SEED)
torch.manual_seed(SEED)
device = "cpu"
if torch.cuda.is_available():
torch.cuda.manual_seed_all(SEED)
device = "cuda"
device'cpu'Dataset¶
Create Dataset¶
For our supervised task, we have a dataset denoted:
where is the number of samples in our dataset.
We model the data generating process under the standard assumption that sample pairs are independent and identically distributed (i.i.d.) according to an unknown joint distribution .
The input data can be represented as a vector:
and the target data can be also represented as a vector:
import random
from sklearn.datasets import make_regression
N: int = 1_000 # number of samples
X, Y = make_regression( # type: ignore
n_samples=N,
n_features=1,
n_targets=1,
bias=random.randint(-5, 5), # random true bias
noise=2,
)
X.shape, Y.shape((1000, 1), (1000,))Y = Y.reshape(-1, 1) # add the axis of length 1
X = X.astype(np.float32)
Y = Y.astype(np.float32)
X.shape, Y.shape((1000, 1), (1000, 1))Split Dataset¶
We partition into three pairwise-disjoint subsets:
Train set : Used to optimize model parameters.
Validation set : Used to tune hyperparameters, monitor over-fitting, and perform model selection.
Test set : Reserved strictly for evaluating unbiased generalization performance on unseen data.
Remark: , and are disjoint, .
Let denote the training input data, the training target data, the validation input data, the validation target data, the test input data, and the target data respectively.
Here we use train_test_split from sklearn to
split our arrays into random train and a test subsets.
from sklearn.model_selection import train_test_split
X_train, X_valid, Y_train, Y_valid = train_test_split(
X,
Y,
test_size=0.2,
random_state=42,
shuffle=True,
)X_train.shape, Y_train.shape((800, 1), (800, 1))X_valid.shape, Y_valid.shape((200, 1), (200, 1))Remark: We have omitted the use of the test set for the rest of this notebook, as our aim here is to understand how the Perceptron works, not to draw general conclusions about a synthetic dataset. We will use the validation set instead of the test set, as it serves the same purpose in this notebook.
Tensor Dataset¶
We convert our NumPy arrays into PyTorch tensors and
wrap them in a TensorDataset to allow for structured iteration.
from collections.abc import Sized
from torch.utils.data import TensorDataset
train_dataset = TensorDataset(torch.from_numpy(X_train), torch.from_numpy(Y_train))
train_dataset[0] # get the first sample(tensor([0.9430]), tensor([4.9075]))valid_dataset = TensorDataset(torch.from_numpy(X_valid), torch.from_numpy(Y_valid))Data Loader¶
from torch.utils.data import DataLoader
BATCH_SIZE: int = 32
train_loader = DataLoader(
train_dataset,
batch_size=BATCH_SIZE,
shuffle=False, # we want to ensure determinism
pin_memory=True,
drop_last=True,
)
len(train_loader), len(train_dataset) // BATCH_SIZE(25, 25)_ = next(iter(train_loader))
_[0].shape, _[1].shape/home/runner/work/inside-deep-learning/inside-deep-learning/.venv/lib/python3.14/site-packages/torch/utils/data/dataloader.py:752: UserWarning: 'pin_memory' argument is set as true but no accelerator is found, then device pinned memory won't be used.
super().__init__(loader)
(torch.Size([32, 1]), torch.Size([32, 1]))valid_loader = DataLoader(
valid_dataset,
batch_size=BATCH_SIZE,
shuffle=False,
pin_memory=True,
drop_last=False,
)
_ = next(iter(valid_loader))
_[0].shape, _[1].shape(torch.Size([32, 1]), torch.Size([32, 1]))Plot Valid Samples¶
To illustrate the relationship between the input and output, we plot only 1000 points.
from matplotlib import pyplot as plt
def plot_lr(
x: np.ndarray | torch.Tensor,
y: np.ndarray | torch.Tensor,
model: SimpleLR | TorchLR | None = None,
) -> None:
plt.scatter(x, y, marker=".", label="Valid data")
if model is not None:
input_ = torch.tensor([x.min(), x.max()]).to(device)
pred = model(input_)
plt.plot(input_.cpu(), pred.cpu(), "r-", label="Predicted")
plt.grid(True)
plt.xlabel("x")
plt.ylabel("y")
plt.legend()
plt.show()x_val, y_val = valid_dataset[:]
plot_lr(x_val, y_val)
Simple Perceptron from Scratch¶
# Empty Linear Regression class
class SimpleLR:
b: torch.Tensor
w: torch.Tensor
def copy_params(self, torch_layer: nn.modules.linear.Linear) -> None:
pass
def __call__(self, x: torch.Tensor) -> torch.Tensor:
return torch.tensor([])
def mse_loss(self, y_true: torch.Tensor, y_pred: torch.Tensor) -> float:
return -1.0
def evaluate(self, x: torch.Tensor, y_true: torch.Tensor) -> float:
return -1.0
def update(
self, x: torch.Tensor, y_true: torch.Tensor, y_pred: torch.Tensor, lr: float
) -> None:
pass
def fit(
self,
train_: torch.utils.data.dataloader.DataLoader,
epochs: int,
lr: float,
valid_: torch.utils.data.dataloader.DataLoader,
) -> None:
passBias and Weight¶
Our model has two trainable parameters , which are called bias and weight respectively.
def init_params(self: SimpleLR) -> None:
self.b = torch.randn(()).to(device) # scalar
self.w = torch.randn(()).to(device) # scalar
setattr(SimpleLR, "__init__", init_params)Let’s add this method copy_params so we can copy the parameter values from a PyTorch class to our scratch class.
def copy_params(self: SimpleLR, torch_layer: nn.modules.linear.Linear) -> None:
"""
Copy the parameters from a module.linear to this model.
Args:
torch_layer: Pytorch module from which to copy the parameters.
"""
self.b.copy_(torch_layer.bias.detach()[0])
self.w.copy_(torch_layer.weight.detach()[0, 0])
setattr(SimpleLR, "copy_params", copy_params)Weighted Sum¶
We selected the weighted sum for as a linear approximation of the true function :
where represents an arbitrary input feature (not necessarily from the training distribution).
To improve computational efficiency, we vectorized calculations in minibatches of data. Given a minibatch of samples:
Note: Let denote the minibatch predicted data of samples.
Remark: Adding a scalar bias to a feature vector relies on tensor broadcasting semantics, implicitly expanding to .
def predict(self: SimpleLR, x: torch.Tensor) -> torch.Tensor:
"""
Predict the output for input x.
Args:
x: Input tensor of shape (m_samples,).
Returns:
y_pred: Predicted output tensor of shape (m_samples,).
"""
return self.b + self.w * x
setattr(SimpleLR, "__call__", predict)We visualize the regression line generated by randomly initialized parameters prior to training:
dummy_model = SimpleLR()
plot_lr(x_val, y_val, dummy_model)
MSE¶
We need a loss function to help guide the adjustment of our parameters during training. We will use Mean Squared Error (MSE) as loss function:
MSE is defined as:
or using a vectorized form:
where is the Euclidean norm or is also called norm (L2 norm).
def mse_loss(self: SimpleLR, y_true: torch.Tensor, y_pred: torch.Tensor) -> float:
"""
MSE loss function between target y_true and y_pred.
Args:
y_true: Target tensor of shape (m_samples,).
y_pred: Predicted tensor of shape (m_samples,).
Returns:
loss: MSE loss between predictions and true values.
"""
return ((y_pred - y_true) ** 2).mean().item()
setattr(SimpleLR, "mse_loss", mse_loss)def evaluate(self: SimpleLR, x: torch.Tensor, y_true: torch.Tensor) -> float:
"""
Evaluate the model on input x and target y_true using MSE.
Args:
x: Input tensor of shape (m_samples,).
y_true: Target tensor of shape (m_samples,).
Returns:
loss: MSE loss between predictions and true values.
"""
y_pred = self(x)
return self.mse_loss(y_true, y_pred)
setattr(SimpleLR, "evaluate", evaluate)Gradients¶
To make our model’s parameters update, it is necessary to compute derivatives.
First, determine the derivatives to be computed
Then, ascertain the shape of each derivative
Finally, compute the derivatives
Using chain rule, we can determine the derivatives we need. Gradient of MSE with respect to bias is:
Gradient of MSE with respect to weight is:
Having defined the necessary derivatives, we now compute their shapes:
MSE Derivative¶
The vectorized form is:
Weighted Sum Derivative¶
Then, the vectorized form is:
Full Chain Rule¶
Derivative of MSE with respect to bias is:
Derivative of MSE with respect to weight is:
Final Gradients¶
Parameters Update¶
Now, let’s update the trainable parameters using gradient descent (GD) as follows:
where is called learning rate.
@torch.inference_mode()
def update(
self: SimpleLR,
x: torch.Tensor,
y_true: torch.Tensor,
y_pred: torch.Tensor,
lr: float,
) -> None:
"""
Update the model parameters.
Args:
x: Input tensor of shape (m_samples,).
y_true: Target tensor of shape (m_samples,).
y_pred: Predicted output tensor of shape (m_samples,).
lr: Learning rate.
"""
delta = 2 * (y_pred - y_true) / len(y_true)
self.b -= lr * delta.sum()
self.w -= lr * torch.matmul(delta.T, x).squeeze()
setattr(SimpleLR, "update", update)Gradient Descent¶
We will use minibatch gradient descent (minibatch GD) to adjust the parameters of our model:
where:
is the number of epochs.
is an arbitrary model’s parameter, in our case are and .
is the number of samples per minibatch.
and are the -th to -th train samples.
Note: are called hyperparameters, because they are adjusted by the developer rather than the model.
Remark: We intentionally do not shuffle the training data at each epoch for the sake of reproducibility. However, in practice, it is recommended to shuffle the training data at each epoch to improve convergence.
To learn more about types of gradient descents, please see gradient descents.
def fit(
self: SimpleLR,
train_: DataLoader,
epochs: int,
lr: float,
valid_: DataLoader,
) -> None:
"""
Fit the model using gradient descent.
Args:
train_: Training dataloader with tensors of shape (m_samples,).
epochs: Number of epochs to fit.
lr: learning rate.
valid_: Validation dataloader with tensors of shape (m_valid_samples,).
"""
# disable lint error: expected "Sized"
assert isinstance(train_.dataset, Sized)
assert isinstance(valid_.dataset, Sized)
for epoch in range(epochs):
# training epoch
running_loss = 0.0
for batch_x, batch_y in train_:
# move only the current batch to VRAM
batch_x = batch_x.to(device, non_blocking=True)
batch_y = batch_y.to(device, non_blocking=True)
# make predictions
y_pred = self(batch_x)
running_loss += self.mse_loss(batch_y, y_pred) * batch_y.size(0)
self.update(batch_x, batch_y, y_pred, lr)
avg_loss = running_loss / len(train_.dataset)
# validation epoch
running_vloss = 0.0
# disable gradient computation and reduce memory consumption.
with torch.no_grad():
for vbatch_x, vbatch_y in valid_:
vbatch_x = vbatch_x.to(device, non_blocking=True)
vbatch_y = vbatch_y.to(device, non_blocking=True)
vy_pred = self(vbatch_x)
running_vloss += self.mse_loss(vbatch_y, vy_pred) * vbatch_y.size(0)
avg_loss_v = running_vloss / len(valid_.dataset)
print(f"epoch: {epoch} - MSE: {avg_loss:.4f} - vMSE: {avg_loss_v:.4f}")
setattr(SimpleLR, "fit", fit)Scratch vs Torch.nn¶
We will be implementing a model created with PyTorch’s pre-built classes for linear regression. This will allow us to compare our model from scratch with the PyTorch model.
PyTorch Module¶
from typing import castclass TorchLR(nn.Module):
def __init__(self, n_features: int) -> None:
super().__init__()
self.layer = nn.Linear(n_features, 1)
self.loss = nn.MSELoss()
def forward(self, x: torch.Tensor) -> torch.Tensor:
return cast(torch.Tensor, self.layer(x))
@torch.inference_mode()
def evaluate(
self,
x: torch.Tensor,
y: torch.Tensor,
) -> float:
self.eval()
y_pred = self(x)
return float(self.loss(y_pred, y).item())
def fit(
self,
train_: DataLoader,
epochs: int,
lr: float,
valid_: DataLoader,
) -> None:
optimizer = torch.optim.SGD(self.parameters(), lr=lr)
# disable lint error: expected "Sized"
assert isinstance(train_.dataset, Sized)
assert isinstance(valid_.dataset, Sized)
for epoch in range(epochs):
self.train() # use model.train() when training
running_loss = 0.0
for batch_x, batch_y in train_:
batch_x = batch_x.to(device, non_blocking=True).unsqueeze(-1)
batch_y = batch_y.to(device, non_blocking=True).unsqueeze(-1)
y_pred = self(batch_x)
loss = self.loss(y_pred, batch_y)
optimizer.zero_grad()
loss.backward()
optimizer.step()
running_loss += loss.item() * batch_y.size(0)
avg_loss = running_loss / len(train_.dataset)
self.eval() # set the model to evaluation mode
with torch.no_grad():
running_vloss = 0.0
for vbatch_x, vbatch_y in valid_:
vbatch_x = vbatch_x.to(device, non_blocking=True)
vbatch_y = vbatch_y.to(device, non_blocking=True)
vy_pred = self(vbatch_x)
vloss = self.loss(vy_pred, vbatch_y)
running_vloss += vloss.item() * vbatch_y.size(0)
avg_loss_v = running_vloss / len(valid_.dataset)
print(f"epoch: {epoch} - MSE: {avg_loss:.4f} - vMSE: {avg_loss_v:.4f}")torch_model = TorchLR(1).to(device)# Let initialize our model from our class
model = SimpleLR()plot_lr(x_val, y_val, model)
Eval¶
We use a norm between PyTorch model and our Scratch model as parameter discrepancy.
def l2(
pred: int | float | torch.Tensor,
true_: int | float | torch.Tensor,
) -> float:
if isinstance(pred, (float, int)) and isinstance(true_, (float, int)):
return abs(pred - true_)
return float(torch.linalg.norm(pred - true_).item())Predictions pre-copy¶
x_val = x_val.to(device, non_blocking=True)
y_val = y_val.to(device, non_blocking=True)We measure the prediction discrepancy between both models using the Euclidean distance:
l2(model(x_val), torch_model(x_val))8.669129371643066The predictions diverge significantly due to independent pseudo-random weight and bias initializations.
Copy Parameters¶
We copy the values of the PyTorch model parameters to our model.
model.copy_params(torch_model.layer)Predictions post-copy¶
We measure the difference between the predictions of both models again.
l2(model(x_val), torch_model(x_val))0.0Discrepancies on the order of fall within the expected machine epsilon () for IEEE 754 single-precision floating-point arithmetic (float32), confirming numerical identity up to hardware precision limits.
Loss¶
l2(model.evaluate(x_val, y_val), torch_model.evaluate(x_val, y_val))0.0Training¶
We are going to train both models using the same hyperparameter values. Given identical initial weights, mini-batch sequences, and learning rates, our analytical gradient updates should produce an optimization trajectory numerically identical to PyTorch’s autograd engine.
LR: float = 0.01 # learning rate
EPOCHS: int = 16 # number of epochsRemark: The training set can be divided evenly into four mini-batches, with each mini-batch containing exactly the same number of examples. Usually, in each training iteration, the training set is shuffled before being divided into mini-batches, and any remaining examples are discarded during that iteration.
model.fit(train_loader, EPOCHS, LR, valid_loader)epoch: 0 - MSE: 48.7781 - vMSE: 27.4725
epoch: 1 - MSE: 19.7109 - vMSE: 12.8196
epoch: 2 - MSE: 9.5582 - vMSE: 7.6048
epoch: 3 - MSE: 6.0183 - vMSE: 5.7289
epoch: 4 - MSE: 4.7881 - vMSE: 5.0425
epoch: 5 - MSE: 4.3629 - vMSE: 4.7845
epoch: 6 - MSE: 4.2174 - vMSE: 4.6837
epoch: 7 - MSE: 4.1684 - vMSE: 4.6421
epoch: 8 - MSE: 4.1525 - vMSE: 4.6237
epoch: 9 - MSE: 4.1476 - vMSE: 4.6150
epoch: 10 - MSE: 4.1464 - vMSE: 4.6106
epoch: 11 - MSE: 4.1462 - vMSE: 4.6083
epoch: 12 - MSE: 4.1462 - vMSE: 4.6070
epoch: 13 - MSE: 4.1464 - vMSE: 4.6062
epoch: 14 - MSE: 4.1465 - vMSE: 4.6058
epoch: 15 - MSE: 4.1465 - vMSE: 4.6055
/home/runner/work/inside-deep-learning/inside-deep-learning/.venv/lib/python3.14/site-packages/torch/utils/data/dataloader.py:752: UserWarning: 'pin_memory' argument is set as true but no accelerator is found, then device pinned memory won't be used.
super().__init__(loader)
torch_model.fit(train_loader, EPOCHS, LR, valid_loader)epoch: 0 - MSE: 48.7781 - vMSE: 27.4725
epoch: 1 - MSE: 19.7109 - vMSE: 12.8196
epoch: 2 - MSE: 9.5582 - vMSE: 7.6048
epoch: 3 - MSE: 6.0183 - vMSE: 5.7289
epoch: 4 - MSE: 4.7881 - vMSE: 5.0425
epoch: 5 - MSE: 4.3629 - vMSE: 4.7845
epoch: 6 - MSE: 4.2174 - vMSE: 4.6837
epoch: 7 - MSE: 4.1684 - vMSE: 4.6421
epoch: 8 - MSE: 4.1525 - vMSE: 4.6237
epoch: 9 - MSE: 4.1476 - vMSE: 4.6150
epoch: 10 - MSE: 4.1464 - vMSE: 4.6106
epoch: 11 - MSE: 4.1462 - vMSE: 4.6083
epoch: 12 - MSE: 4.1462 - vMSE: 4.6070
epoch: 13 - MSE: 4.1464 - vMSE: 4.6062
epoch: 14 - MSE: 4.1465 - vMSE: 4.6058
epoch: 15 - MSE: 4.1465 - vMSE: 4.6055
/home/runner/work/inside-deep-learning/inside-deep-learning/.venv/lib/python3.14/site-packages/torch/utils/data/dataloader.py:752: UserWarning: 'pin_memory' argument is set as true but no accelerator is found, then device pinned memory won't be used.
super().__init__(loader)
Predictions after Training¶
l2(model(x_val), torch_model(x_val))0.0plot_lr(x_val.cpu(), y_val.cpu(), model)
Bias Comparison¶
We directly measure the difference between the bias values of both models.
l2(model.b.clone(), torch_model.layer.bias.detach()[0])0.0Weight Comparison¶
And measure the difference between the weight values of both models.
l2(model.w.clone(), torch_model.layer.weight.detach()[0, 0])0.0Conclusion¶
In this notebook, we dissected the simplest foundational block of deep learning: the single-variable linear perceptron trained via Ordinary Least Squares (OLS) under Mean Squared Error (MSE).
Key Takeaways:¶
First-Principles Derivation: By applying the multivariable chain rule alongside Kronecker delta identities (), we derived exact analytical expressions for parameter gradients ( and ). We demonstrated that expressing these derivatives in vectorized form () eliminates scalar loops and maps directly to hardware-accelerated tensor contractions.
Numerical Equivalence with Autodiff: Benchmarking our custom implementation against PyTorch’s native
nn.Linearand automatic differentiation engine yielded parameter and prediction discrepancies on the order of 10-7. This proves that our manual gradient updates match PyTorch’s backward graph computation within the bounds of IEEE 754 single-precision floating-point tolerance.Optimization Dynamics: Mini-batch gradient descent provided stable, monotonically decreasing loss trajectories across both training and validation sets, verifying that our batch-averaged gradient scaling correctly stabilizes the step size regardless of batch cardinality.
Next Steps:¶
Multivariate Linear Regression (Chapter 1.2): Extending input features from a scalar to a -dimensional feature vector , requiring weight vectors and full matrix operations ().

