DaZu / PyTorchDaZu home

4 · Text / Quick guide

Tokens and text classifiers

Represent sentences as tensors, batch different lengths, and compare pooled embeddings with contextual models.

The input changes from image pixels to token IDs, but the logits, loss, optimizer and validation logic remain the same.

How it works

Wordsclear image / bad blurry photolengths 2 and 3
Toy vocabularyclear=2, image=3, bad=4, blurry=5, photo=6IDs are keys, not numeric meaning
Padded batch[2,3,0] / [4,5,6]mask [1,1,0] / [1,1,1]
Embedding lookup[2,3,E]one vector per token position
Masked mean[2,E]divide each sum by its real token count

Toy vocabulary illustration. The runnable project has its own saved training-only vocabulary; pretrained tokenizer IDs depend on the checkpoint.

offset 0 → 2 3offset 2 → 4 5 6

Five concatenated IDs, two starts, two pooled vectors. Reversing tokens within either group leaves mean pooling unchanged.

Tokenization breaks text into words, subwords or characters. A vocabulary maps tokens to integer IDs; embeddings map those IDs to vectors. A classifier turns the resulting representation into class scores.

TokenizeText → words or subwordsBatchIDs + offsets or padding maskRepresentEmbeddings → pooled/context vectorClassifyLogits → loss → validation

An ID is a lookup key, not a measure of meaning. Static embeddings give a token one vector; contextual models use surrounding tokens, so “bat” can have a different representation in a sports sentence and an animal sentence.

Code and functions

Function / settingWhat it doesWhat to check
nn.EmbeddingLook up one trainable vector per token ID.IDs use long; keep padding and unknown IDs distinct.
EmbeddingBag / offsetsPool concatenated variable-length sequences without padding.Offsets mark each sentence start. Mean pooling loses word order.
pad_sequence / attention_maskBatch different lengths and identify real tokens.A manual mean divides by the real token count, not the padded length.
DataCollatorWithPaddingPad a pretrained-tokenizer batch to its longest input.Truncation bounds cost but can remove important words.
CrossEntropyLoss(weight=...)Give classes different loss priorities.Derive weights from training labels and align them with class IDs.
DistilBERT / .logitsUse a contextual model with a classification head.Use its matching tokenizer; a newly added head needs training.
Two variable-length sentences · no padding required
ids = torch.tensor([2, 3, 4, 5, 6], dtype=torch.long)
offsets = torch.tensor([0, 2], dtype=torch.long)
embedding = torch.nn.EmbeddingBag(7, 8, mode="mean")
vectors = embedding(ids, offsets)  # [2, 8]
head = torch.nn.Linear(8, 2)
logits = head(vectors)             # [2, 2]

Check your understanding

The example pools [2,3] and [4,5,6] into two vectors. A third sentence needs one more start offset, not padding to the longest sentence.

A pooled baseline is cheap and inspectable, but “dog bites person” and “person bites dog” have the same bag of words. Use an order-aware contextual model when that distinction matters.

Fine-tuning: compare a new head alone, the head plus selected late blocks, or all layers. Build the optimizer after deciding what is trainable. Dynamic padding reduces unused positions; it does not eliminate every padded computation.

Common mistakes