DaZu / PyTorchDaZu home

2 · Training / Quick guide

Metrics, tuning and efficiency

Choose meaningful metrics, compare training settings fairly, and reduce time or memory where the run needs it.

Builds on the Fundamentals loop. The same experiment controls apply to the Vision and Text models that follow.

How it works

Inspection outcomes → metrics

Positive means “defect”. Rows are truth; columns are predictions.

Exact arithmetic on editable counts; no trained detector.
Microbatch 116 samples → backwardgradients accumulate; weights stay fixed
Microbatch 216 samples → backwardadd to existing gradients
Microbatch 316 samples → backwardaverage by 48 actual samples
Updateoptimizer.step()one update, then clear gradients

Single-device illustration. For 133 samples, full groups are 48 and 48; the final update uses 37.

Keep the basic training loop. First define a validation measure, then compare settings on the same data split and budget. Only optimize speed after you know whether the model is learning.

MeasureLoss and class-level errorsTuneLearning rate and regularizationProfileLoading, compute and memoryCompareQuality plus time and cost

Accuracy can hide rare-class failures. Precision measures how many predicted positives are correct; recall measures how many actual positives were found. Macro F1 gives each class equal weight. Keep the final test set outside tuning.

Code and functions

Function / settingWhat it doesWhat to check
lr / weight_decay / batch_sizeSet update size, regularization and samples per batch.Change a small number of settings; reuse the same split and starting weights.
StepLR / CosineAnnealingLRReduce learning rate on a chosen schedule.A scheduler call advances its clock; match calls to epochs or steps intentionally.
ReduceLROnPlateauReduce LR when a validation measure stalls.After validation: loss needs mode="min"; a quality score needs mode="max".
Optuna trial / studySearch configurations using a validation objective.Each trial needs a fresh model; the objective must reflect your actual goal.
num_workers / pin_memoryPrepare batches and support host-to-CUDA transfer.Start with zero workers for debugging; more workers can use more RAM or be slower.
autocast / GradScalerUse mixed precision for suitable GPU operations.Check device/dtype support; FP16 scaling and clipping order matter.
Trainer / profilerLightning orchestrates training; a profiler locates expensive steps.Do not mix manual optimizer updates with automatic ones. Benchmark without profiling overhead.
Epoch-level plateau scheduler · after training and validation
scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(
    optimizer, mode="min", factor=0.5, patience=2
)
# In the epoch loop, after val_loss has been computed:
scheduler.step(val_loss)

Check your understanding

Accumulation: a microbatch of 32 accumulated for four steps gives 128 samples per full optimizer update on one device. It avoids keeping four batches of activations alive at once. Divide gradients by the actual sample count and flush a short final group.

Accumulation does not make BatchNorm see 128 samples at once. Its statistics still come from each microbatch.

Speed: if input loading dominates, a faster GPU computation alone will have little effect. Measure loading and model work separately. Synchronize CUDA around timings and include a warmup.

Recall check: would you choose higher accuracy if minority-class recall falls? Decide which errors matter before selecting the winning run.

Common mistakes