24/08/2026
𝗚𝗿𝗮𝗱𝗶𝗲𝗻𝘁 𝗗𝗲𝘀𝗰𝗲𝗻𝘁 is an optimization algorithm used in 𝗠𝗮𝗰𝗵𝗶𝗻𝗲 𝗟𝗲𝗮𝗿𝗻𝗶𝗻𝗴 to find the values of model parameters (weights and biases) that minimize a 𝗹𝗼𝘀𝘀/𝗰𝗼𝘀𝘁 𝗳𝘂𝗻𝗰𝘁𝗶𝗼𝗻.
🔹 𝗦𝗶𝗺𝗽𝗹𝗲 𝗶𝗱𝗲𝗮
Imagine you are standing on a mountain and want to reach the 𝗹𝗼𝘄𝗲𝘀𝘁 𝗽𝗼𝗶𝗻𝘁 in the valley, but you cannot see the entire landscape.
You can:
1️⃣ Look at the 𝘀𝗹𝗼𝗽𝗲 around you.
2️⃣ Determine which direction goes downhill.
3️⃣ Take a small step in that direction.
4️⃣ Repeat until you reach a low point.
That is essentially 𝗚𝗿𝗮𝗱𝗶𝗲𝗻𝘁 𝗗𝗲𝘀𝗰𝗲𝗻𝘁.
🔹 𝗜𝗻 𝗠𝗮𝗰𝗵𝗶𝗻𝗲 𝗟𝗲𝗮𝗿𝗻𝗶𝗻𝗴
Suppose our model has a parameter 𝘄, and its loss is 𝗝(𝘄).
The update rule is:
𝘄ₙₑw = 𝘄ₒₗd − α × ∂𝗝/∂𝘄
where:
• 𝘄 = model parameter/weight
• 𝗝 = loss function
• ∂𝗝/∂𝘄 = gradient, telling us the direction of steepest increase
• α = learning rate, controlling how big a step we take
The 𝗺𝗶𝗻𝘂𝘀 𝘀𝗶𝗴𝗻 is important: we move in the direction opposite to the gradient because we want to 𝗿𝗲𝗱𝘂𝗰𝗲 𝘁𝗵𝗲 𝗹𝗼𝘀𝘀.
🔹 𝗘𝘅𝗮𝗺𝗽𝗹𝗲
Suppose:
• Current weight = 𝟱
• Gradient = 𝟮
• Learning rate = 𝟬.𝟭
Then:
𝘄ₙₑw = 𝟱 − (𝟬.𝟭 × 𝟮) = 𝟰.𝟴
So the weight changes from 𝟱 → 𝟰.𝟴.
On the next iteration, the gradient is recalculated and the process continues.
🔹 𝗪𝗵𝘆 𝗶𝘀 𝗶𝘁 𝗶𝗺𝗽𝗼𝗿𝘁𝗮𝗻𝘁?
Training a neural network essentially involves repeating:
𝗣𝗿𝗲𝗱𝗶𝗰𝘁𝗶𝗼𝗻 → 𝗖𝗮𝗹𝗰𝘂𝗹𝗮𝘁𝗲 𝗟𝗼𝘀𝘀 → 𝗖𝗮𝗹𝗰𝘂𝗹𝗮𝘁𝗲 𝗚𝗿𝗮𝗱𝗶𝗲𝗻𝘁𝘀 → 𝗨𝗽𝗱𝗮𝘁𝗲 𝗪𝗲𝗶𝗴𝗵𝘁𝘀
Input
↓
Neural Network
↓
Prediction
↓
Loss Function
↓
Gradient
↓
Update Weights
↓
Repeat ↺
Eventually, the model ideally reaches a set of weights where the loss is very small.
👉 𝗜𝗻 𝗼𝗻𝗲 𝘀𝗲𝗻𝘁𝗲𝗻𝗰𝗲:
𝗚𝗿𝗮𝗱𝗶𝗲𝗻𝘁 𝗗𝗲𝘀𝗰𝗲𝗻𝘁 𝗶𝘀 𝗮 𝗺𝗲𝘁𝗵𝗼𝗱 𝗳𝗼𝗿 𝗴𝗿𝗮𝗱𝘂𝗮𝗹𝗹𝘆 𝗮𝗱𝗷𝘂𝘀𝘁𝗶𝗻𝗴 𝗮 𝗺𝗼𝗱𝗲𝗹'𝘀 𝗽𝗮𝗿𝗮𝗺𝗲𝘁𝗲𝗿𝘀 𝗶𝗻 𝘁𝗵𝗲 𝗱𝗶𝗿𝗲𝗰𝘁𝗶𝗼𝗻 𝘁𝗵𝗮𝘁 𝗿𝗲𝗱𝘂𝗰𝗲𝘀 𝗶𝘁𝘀 𝗽𝗿𝗲𝗱𝗶𝗰𝘁𝗶𝗼𝗻 𝗲𝗿𝗿𝗼𝗿.