DaZu / PyTorchDaZu home

3 · Vision / Quick guide

Images and pretrained models

Prepare images consistently, use controlled augmentation and noise, and adapt pretrained features to your classes.

Uses the same loop from Fundamentals and the validation rules from Training. Here the extra decisions are image preprocessing, output type and which layers to train.

How it works

Illustrative scratched panel classified as scratched
Classification: one class for the image · [N,K]
Illustrative panel with boxes locating its two scratches
Detection: boxes, labels and scores per image
Illustrative panel with colored masks covering scratch pixels
Segmentation: class per pixel · [N,K,H,W]

Original illustrative annotations, not predictions from a trained model.

Head onlybackbone frozen / new head learnskeep backbone BatchNorm statistics fixed
Partial fine-tuninglate block + head learninclude newly unfrozen parameters in optimizer
Full fine-tuningall pretrained layers learnhigher backward and optimizer cost

Illustration: shapes and operations, not measured model performance.

A vision pipeline reads an image, applies training variation, converts it to a tensor, and normalizes it for the chosen model. Validation uses stable preprocessing so changing augmentations do not change the comparison.

ReadRGB image and known labelVaryTraining-only crop, colour, noisePrepareFloat tensor → normalizePredictScores, boxes or pixel labels

Classification returns [N,K] scores; detection returns boxes, labels and scores per image; segmentation returns [N,K,H,W] pixel scores. Drawing a box or mask displays an existing prediction.

Code and functions

Function / settingWhat it doesWhat to check
ImageFolderBuild a dataset from class folders.Save the class-to-ID map; train/validation folder sets must agree.
ToTensor / NormalizeConvert typical byte images to 0–1 floats, then standardize channels.Tensor shape is [C,H,W]. Negative normalized values are expected.
RandomResizedCrop / ColorJitterAdd geometry or colour variation.Inspect samples; the transformed image must still justify its original label.
weights.transforms()Use preprocessing associated with pretrained weights.Keep it with the model for validation and inference.
requires_grad_(False)Stop gradients for selected parameters.It does not freeze BatchNorm statistics or eliminate backbone forward computation.
model.fc = nn.Linear(...)Replace a ResNet classification head.Use your class count and optimize the new head; other architectures use other paths.
ResNet head-only setup · weights may download on first use
from torchvision.models import resnet18, ResNet18_Weights

weights = ResNet18_Weights.DEFAULT
model = resnet18(weights=weights)
model.requires_grad_(False)
model.fc = torch.nn.Linear(model.fc.in_features, num_classes)
optimizer = torch.optim.Adam(model.fc.parameters(), lr=1e-3)
# For head-only training with fixed backbone statistics:
model.eval()
model.fc.train()

Start by training the head. If validation warrants it, unfreeze selected late blocks at a smaller learning rate and rebuild the optimizer. Full fine-tuning updates every layer.

Check your understanding

Salt and pepper noise replaces selected pixels with black or white. This is the custom noise used in the course. Gaussian noise adds fluctuations; camera shot noise depends on signal intensity. These are different mechanisms; a simple transform only approximates part of camera behaviour.

For a float-tensor noise transform: convert to 0–1 → add noise → normalize. Keep ordinary validation clean and deterministic. A separate noisy validation set tests robustness.

Recall check: would you horizontally flip text or rotate a “6” into a “9”? An augmentation is only useful when the label remains valid.

Common mistakes