As the final step in developing the model, rather than modelling the relationship between multiple inputs and a single output, we will have multiple outputs at the same time. This problem is known as Multioutput Linear Regression.
Figure 1:Multioutput perceptron with multiple inputs and multiple outputs.
We can think of this model as a single neuron that takes multiple inputs and returns multiple outputs; we can also think of it as a dense layer containing multiple neurons.
Figure 2:Multioutput perceptron as a dense layer.
Both interpretations are valid; the idea of having a single neuron that returns multiple outputs is a useful concept for the following section Classification, when we introduce the activation function. The concept of grouping neurons into a single layer is useful for chapter Multilayer Perceptron, when we introduce multiple dense layers.
We assume that the true unknown function that maps the relationship between the input and output is:
where is an intrinsic noise independent of , normally . Note that the target is now a vector.
The goal of multioutput linear regression is estimate to by a linear approximation such that:
This means that we want to estimate multiple output features from multiple input features.
Purpose of this Notebook:
Create a dataset for multioutput linear regression task
Create our own Multioutput Perceptron class from scratch
Calculate the gradients from scratch
Implement gradient descent from scratch
Train our Perceptron
Compare our Perceptron to PyTorch’s built-in implementation
Setup¶
print('Start package installation...')Start package installation...
%%capture
%pip install torch
%pip install scikit-learnprint('Packages installed successfully!')Packages installed successfully!
import torch
from torch import nn
from platform import python_version
python_version(), torch.__version__('3.14.7', '2.12.0+cpu')# Set seeds for reproducibility
device = 'cpu'
torch.manual_seed(13)
if torch.cuda.is_available():
torch.cuda.manual_seed_all(13)
device = 'cuda'
import numpy as np; np.random.seed(13)
import random; random.seed(13)
device'cpu'The add_to_class function is used to add new methods to a previously defined class; we do this to gradually enhance the class.
def add_to_class(Class):
"""Register functions as methods in created class."""
def wrapper(obj): setattr(Class, obj.__name__, obj)
return wrapperDataset¶
Create Dataset¶
The dataset consists of input-target pairs :
where denotes the -th input sample with input features and its corresponding target value with output features.
The target data also can be represented as a matrix :
from sklearn.datasets import make_regression
import random
N: int = 1_500 # number of samples
D: int = 5 # number of input features
C: int = 3 # number of output features
X, Y = make_regression( # type: ignore
n_samples=N,
n_features=D,
n_targets=C,
n_informative=N - 1,
bias=random.random(),
noise=1
)
X = X.astype(np.float32)
Y = Y.astype(np.float32)
X.shape, Y.shape---------------------------------------------------------------------------
NameError Traceback (most recent call last)
Cell In[7], line 17
13 bias=random.random(),
14 noise=1
15 )
16
---> 17 X = X.astype(np.float32)
18 Y = Y.astype(np.float32)
19
20 X.shape, Y.shape
NameError: name 'np' is not definedSplit Dataset¶
from sklearn.model_selection import train_test_split
X_train, X_valid, Y_train, Y_valid = train_test_split(
X, Y,
test_size=0.2,
random_state=42,
shuffle=True,
)X_train.shape, Y_train.shapeX_valid.shape, Y_valid.shapeRemark: We have omitted the use of the test set for the rest of this notebook, as our aim here is to understand how the Perceptron works, not to draw general conclusions about a synthetic dataset. We will use the validation set instead of the test set, as it serves the same purpose in this notebook.
Tensor Dataset¶
from torch.utils.data import TensorDatasettrain_dataset = TensorDataset(
torch.from_numpy(X_train),
torch.from_numpy(Y_train)
)Get the first train input sample:
train_dataset[0][0]And get the first target sample:
train_dataset[0][1]valid_dataset = TensorDataset(
torch.from_numpy(X_valid),
torch.from_numpy(Y_valid)
)Data Loader¶
from torch.utils.data import DataLoader
BATCH_SIZE: int = 32train_loader = DataLoader(
train_dataset,
batch_size=BATCH_SIZE,
shuffle=False, # we want to ensure determinism
pin_memory=True,
)valid_loader = DataLoader(
valid_dataset,
batch_size=BATCH_SIZE,
shuffle=False,
pin_memory=True,
)Scratch Multioutput Perceptron¶
Bias and Weight¶
Our model has two trainable parameters called bias and weight respectively
class MultioutputRegression:
def __init__(self, n_features: int, out_features: int):
self.b = torch.randn(out_features).to(device)
self.w = torch.randn(n_features, out_features).to(device)
def copy_params(self, torch_layer: nn.modules.linear.Linear):
"""
Copy the parameters from a module.linear to this model.
Args:
torch_layer: Pytorch module from which to copy the parameters.
"""
self.b.copy_(torch_layer.bias.detach())
self.w.copy_(torch_layer.weight.T.detach())Weighted Sum¶
The weight sum is a vector-matrix multiplication between the input and weight plus a vector summation using broadcasting mechanism:
Notice that the predicted output is a vector.
For vectorization, given a minibatch input of samples and input features, we can compute the weighted sum as:
where the predict matrix of output features.
Note: is the number of samples in the dataset, while is the number of samples in a minibatch.
The prediction for the -th sample and output feature is:
This formula will be useful for gradient descent.
@add_to_class(MultioutputRegression)
def predict(self, x: torch.Tensor) -> torch.Tensor:
"""
Predict the output for input x
Args:
x: Input tensor of shape (m_samples, d_features).
Returns:
y_pred: Predicted output tensor of shape (m_samples, c_features).
"""
return torch.matmul(x, self.w) + self.bMSE¶
The MSE function as loss function needs a little update from its predecessor for multioutput features:
MSE is defined as:
where is the -th sample/row of as a vector , and is the -th column of the weight matrix as a vector .
For vectorization, we can compute MSE as:
where .
Note: is called Hadamard product and performs element-wise product.
@add_to_class(MultioutputRegression)
def mse_loss(self, y_true: torch.Tensor, y_pred: torch.Tensor):
"""
MSE loss function between target y_true and y_pred.
Args:
y_true: Target tensor of shape (m_samples, c_features).
y_pred: Predicted tensor of shape (m_samples, c_features).
Returns:
loss: MSE loss between predictions and true values.
"""
return ((y_pred - y_true) ** 2).mean().item()
@add_to_class(MultioutputRegression)
def evaluate(self, x: torch.Tensor, y_true: torch.Tensor):
"""
Evaluate the model on input x and target y_true using MSE.
Args:
x: Input tensor of shape (m_samples, d_features).
y_true: Target tensor of shape (m_samples, c_features).
Returns:
loss: MSE loss between predictions and true values.
"""
y_pred = self.predict(x)
return self.mse_loss(y_true, y_pred)Gradients¶
Derivative of MSE with respect to bias:
and derivative of MSE with respect to weight:
where the shape of each derivative is:
Note: Third- and fourth-order derivatives may seem complicated, but you’ll soon find that the Kronecker delta property reduces the order when we differentiate.
Remark: The way to calculate derivatives using mixed-layout is to multiply the order of the dependent variable by the order of the independent variable. Orders of 0, such as the loss function, are not included.
For example, let and , then:
The use of brackets is purely for visual purposes and does not alter the order of the derivative.
MSE derivative¶
The vectorized form is:
Weighted Sum Derivative¶
With Respect to Bias¶
for , and .
With Respect to Weight¶
for , for , and .
Full Chain Rule¶
The vectorized form is:
Note: Let be the -th input features from all input in a minibatch:
and let be the -th output feature from all the predicted and target minibatch respectively:
The vectorized form is:
Final Gradients¶
Parameters Update¶
where is the learning rate.
@add_to_class(MultioutputRegression)
def update(self, x: torch.Tensor, y_true: torch.Tensor,
y_pred: torch.Tensor, lr: float):
"""
Update the model parameters.
Args:
x: Input tensor of shape (m_samples, d_features).
y_true: Target tensor of shape (m_samples, c_features).
y_pred: Predicted output tensor of shape (m_samples, c_features).
lr: Learning rate.
"""
delta = 2 * (y_pred - y_true) / y_true.numel()
self.b -= lr * delta.sum(dim=0)
self.w -= lr * torch.matmul(x.T, delta)Gradient Descent¶
@add_to_class(MultioutputRegression)
def fit(self, train_loader: DataLoader,
epochs: int, lr: float,
valid_loader: DataLoader):
"""
Fit the model using gradient descent.
Args:
train_loader: Train dataloader.
epochs: Number of epochs to fit.
lr: Learning rate.
valid_loader: Valid dataloader.
"""
for epoch in range(epochs):
# training epoch
running_loss = 0.0
for batch_x, batch_y in train_loader:
# move only the current batch to VRAM
batch_x = batch_x.to(device, non_blocking=True)
batch_y = batch_y.to(device, non_blocking=True)
# make predictions
y_pred = self.predict(batch_x)
running_loss += self.mse_loss(
batch_y, y_pred
)
self.update(
batch_x, batch_y,
y_pred, lr
)
avg_loss = running_loss / len(train_loader)
# validation epoch
running_vloss = 0.0
# disable gradient computation and reduce memory consumption.
with torch.no_grad():
for vbatch_x, vbatch_y in valid_loader:
vbatch_x = vbatch_x.to(device, non_blocking=True)
vbatch_y = vbatch_y.to(device, non_blocking=True)
vy_pred = self.predict(vbatch_x)
running_vloss += self.mse_loss(
vbatch_y, vy_pred
)
avg_loss_v = running_vloss / len(valid_loader)
print(f'epoch: {epoch} - MSE: {avg_loss:.4f} - vMSE: {avg_loss_v:.4f}')Scratch vs Torch.nn¶
Torch.nn model¶
class TorchLinearRegression(nn.Module):
def __init__(self, d_features, c_out_features):
super().__init__()
self.layer = nn.Linear(
d_features,
c_out_features
)
self.loss = nn.MSELoss()
def forward(self, x):
return self.layer(x)
@torch.inference_mode()
def evaluate(self, x, y):
self.eval()
y_pred = self(x)
return self.loss(y_pred, y).item()
def fit(self, train_loader,
epochs, lr,
valid_loader):
optimizer = torch.optim.SGD(self.parameters(), lr=lr)
for epoch in range(epochs):
self.train() # use model.train() when training
running_loss = 0.0 # train loss
for batch_x, batch_y in train_loader:
batch_x = batch_x.to(device, non_blocking=True)
batch_y = batch_y.to(device, non_blocking=True)
y_pred = self(batch_x)
loss = self.loss(y_pred, batch_y)
optimizer.zero_grad()
loss.backward()
optimizer.step()
running_loss += loss.item()
avg_loss = running_loss / len(train_loader)
# validation epoch
self.eval() # set the model to evaluation mode
with torch.no_grad():
running_vloss = 0.0
for vbatch_x, vbatch_y in valid_loader:
vbatch_x = vbatch_x.to(device, non_blocking=True)
vbatch_y = vbatch_y.to(device, non_blocking=True)
vy_pred = self(vbatch_x)
vloss = self.loss(vy_pred, vbatch_y)
running_vloss += vloss.item()
avg_loss_v = running_vloss / len(valid_loader)
print(f'epoch: {epoch} - MSE: {avg_loss:.4f} - vMSE: {avg_loss_v:.4f}')torch_model = TorchLinearRegression(D, C).to(device)model = MultioutputRegression(D, C)Eval¶
We use a norm between PyTorch model and our Scratch model as parameter discrepancy.
def l2(pred, true_):
if isinstance(pred, (float, int)) and isinstance(true_, (float, int)):
return abs(pred - true_)
return torch.linalg.norm(pred - true_).item()Predictions pre-copy¶
Let compare the predictions of our model and PyTorch implementation using norm:
x_valid, y_valid = valid_dataset[:]
x_valid = x_valid.to(device, non_blocking=True)
y_valid = y_valid.to(device, non_blocking=True)l2(
model.predict(x_valid),
torch_model(x_valid)
)They differ considerably because each model has its own parameters initialized randomly and independently of the other model.
Copy Parameters¶
We copy the values of the PyTorch model parameters to our model.
model.copy_params(torch_model.layer)Predictions post-copy¶
We measure the difference between the predictions of both models again.
l2(
model.predict(x_valid),
torch_model(x_valid)
)Loss¶
l2(
model.evaluate(x_valid, y_valid),
torch_model.evaluate(x_valid, y_valid)
)Training¶
We are going to train both models using the same hyperparameter values. If our model is well designed, then starting from the same parameters it should arrive at the same parameters’ values as the PyTorch model after training.
LR: float = 0.01 # learning rate
EPOCHS: int = 16 # number of epochsmodel.fit(
train_loader,
EPOCHS, LR,
valid_loader
)torch_model.fit(
train_loader,
EPOCHS, LR,
valid_loader
)Predictions after Training¶
l2(
model.predict(x_valid),
torch_model(x_valid)
)Bias Comparison¶
We directly measure the difference between the bias values of both models.
l2(
model.b.clone(),
torch_model.layer.bias.detach()
)Weight Comparison¶
l2(
model.w.clone(),
torch_model.layer.weight.detach().T
)Our model, built from scratch, works in the same way as the PyTorch’s built-in implementation according to . This notebook serves as the basis for the following section, where the weighted sum is not modified, but a new element activation function is added to it in order to model a new type of problem: Classification. That is why it is important to understand how and why the weighted sum works before adding further elements to our models.

