2 · Training / Quick guide
Metrics, tuning and efficiency
Choose meaningful metrics, compare training settings fairly, and reduce time or memory where the run needs it.
Builds on the Fundamentals loop. The same experiment controls apply to the Vision and Text models that follow.
How it works
Positive means “defect”. Rows are truth; columns are predictions.
16 samples → backwardgradients accumulate; weights stay fixed16 samples → backwardadd to existing gradients16 samples → backwardaverage by 48 actual samplesoptimizer.step()one update, then clear gradientsSingle-device illustration. For 133 samples, full groups are 48 and 48; the final update uses 37.
Keep the basic training loop. First define a validation measure, then compare settings on the same data split and budget. Only optimize speed after you know whether the model is learning.
Accuracy can hide rare-class failures. Precision measures how many predicted positives are correct; recall measures how many actual positives were found. Macro F1 gives each class equal weight. Keep the final test set outside tuning.
Code and functions
| Function / setting | What it does | What to check |
|---|---|---|
lr / weight_decay / batch_size | Set update size, regularization and samples per batch. | Change a small number of settings; reuse the same split and starting weights. |
StepLR / CosineAnnealingLR | Reduce learning rate on a chosen schedule. | A scheduler call advances its clock; match calls to epochs or steps intentionally. |
ReduceLROnPlateau | Reduce LR when a validation measure stalls. | After validation: loss needs mode="min"; a quality score needs mode="max". |
Optuna trial / study | Search configurations using a validation objective. | Each trial needs a fresh model; the objective must reflect your actual goal. |
num_workers / pin_memory | Prepare batches and support host-to-CUDA transfer. | Start with zero workers for debugging; more workers can use more RAM or be slower. |
autocast / GradScaler | Use mixed precision for suitable GPU operations. | Check device/dtype support; FP16 scaling and clipping order matter. |
Trainer / profiler | Lightning orchestrates training; a profiler locates expensive steps. | Do not mix manual optimizer updates with automatic ones. Benchmark without profiling overhead. |
scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(
optimizer, mode="min", factor=0.5, patience=2
)
# In the epoch loop, after val_loss has been computed:
scheduler.step(val_loss)Check your understanding
Accumulation: a microbatch of 32 accumulated for four steps gives 128 samples per full optimizer update on one device. It avoids keeping four batches of activations alive at once. Divide gradients by the actual sample count and flush a short final group.
Accumulation does not make BatchNorm see 128 samples at once. Its statistics still come from each microbatch.
Speed: if input loading dominates, a faster GPU computation alone will have little effect. Measure loading and model work separately. Synchronize CUDA around timings and include a warmup.
Recall check: would you choose higher accuracy if minority-class recall falls? Decide which errors matter before selecting the winning run.
Common mistakes
- Comparing settings with different splits, seeds or epoch budgets.
- Averaging batch accuracies equally when batch sizes differ; aggregate sample counts or metric state over the full evaluation.
- Calling a plateau scheduler without its validation measure, or stepping it before validation.
- Increasing workers and prefetching until RAM becomes the new bottleneck.
- Reporting parameter size as peak training memory; activations and optimizer state also consume memory.