Gradient Descent
August 10, 2025 · 1 min read
You begin with random weights. This is the honest starting point. No prior knowledge. No special insight. Just numbers, drawn from a distribution, waiting to be wrong.
The loss is large at first. This is expected. You compute the gradient - the direction of steepest ascent - and step the other way.
Not far. Never too far. The learning rate enforces humility. A large step might cross the valley you were trying to reach. Better to creep.
Epoch after epoch, the loss curves downward. Not smoothly. The noise is real. Plateaus. Local minima. Saddle points that feel like progress until they don't.
But you continue. The objective is patient. It doesn't care how long the training takes, only whether the loss is lower than it was before.
And eventually - not always, but often - convergence. The weights settle. The model has learned what the data had to teach.
It knows nothing it wasn't shown. It forgets nothing it was shown enough times. This is not intelligence. But it is something.
Written by Dhruv Choudhary, AI Engineer
AI Engineer at AI LifeBOT, where I build GenAI systems that ship to production, not just notebooks and demos. Shipping RAG pipelines and agentic systems into real government and healthcare deployments.