Deep learning trains layered parameterized functions from examples. Architecture choice matters, but the data, objective, optimization process, and validation design determine whether the result holds outside a training run.
Choose a starting point
Start with tensors, gradients, and training loops. Then compare network families, representation learning, transfer, evaluation, and serving. Track compute and memory costs with each modeling choice.
- Deep Learning Tutorial
- Tensor contracts: shape, dtype, device and mask
- Training and validation modes: measure the model you will serve
- Gradient accumulation and effective batch accounting
- Scaled dot-product attention and mask contracts
- Convolution output geometry and receptive fields
- Frozen features versus fine-tuning a pretrained backbone
- Packed LSTM sequences and the last valid state
- Masked reconstruction loss and anomaly thresholds
Common Mistakes
Loss can fall while a model overfits, exploits a shortcut, or becomes too expensive to serve. The linked projects and diagnostics help separate these outcomes.
