Building my own autodiff module
Forward and backward passes clicked once I treated every op as a graph node.
For SYDE 577 assignment one we couldn’t use PyTorch. We had to build autodiff ourselves. I sketched a tiny graph first: two parameters, one multiply, one loss. Each node stores its value, accumulates gradients from children, and implements forward() and backward(grad).
Matrix multiply’s backward pass took the longest. Grad w.r.t. both W and x. Once that worked, stacking activations and a simple trainer was mechanical. PyTorch hides a lot of bookkeeping.