4 · Text / Quick guide
Tokens and text classifiers
Represent sentences as tensors, batch different lengths, and compare pooled embeddings with contextual models.
The input changes from image pixels to token IDs, but the logits, loss, optimizer and validation logic remain the same.
How it works
clear image / bad blurry photolengths 2 and 3clear=2, image=3, bad=4, blurry=5, photo=6IDs are keys, not numeric meaning[2,3,0] / [4,5,6]mask [1,1,0] / [1,1,1][2,3,E]one vector per token position[2,E]divide each sum by its real token countToy vocabulary illustration. The runnable project has its own saved training-only vocabulary; pretrained tokenizer IDs depend on the checkpoint.
Five concatenated IDs, two starts, two pooled vectors. Reversing tokens within either group leaves mean pooling unchanged.
Tokenization breaks text into words, subwords or characters. A vocabulary maps tokens to integer IDs; embeddings map those IDs to vectors. A classifier turns the resulting representation into class scores.
An ID is a lookup key, not a measure of meaning. Static embeddings give a token one vector; contextual models use surrounding tokens, so “bat” can have a different representation in a sports sentence and an animal sentence.
Code and functions
| Function / setting | What it does | What to check |
|---|---|---|
nn.Embedding | Look up one trainable vector per token ID. | IDs use long; keep padding and unknown IDs distinct. |
EmbeddingBag / offsets | Pool concatenated variable-length sequences without padding. | Offsets mark each sentence start. Mean pooling loses word order. |
pad_sequence / attention_mask | Batch different lengths and identify real tokens. | A manual mean divides by the real token count, not the padded length. |
DataCollatorWithPadding | Pad a pretrained-tokenizer batch to its longest input. | Truncation bounds cost but can remove important words. |
CrossEntropyLoss(weight=...) | Give classes different loss priorities. | Derive weights from training labels and align them with class IDs. |
DistilBERT / .logits | Use a contextual model with a classification head. | Use its matching tokenizer; a newly added head needs training. |
ids = torch.tensor([2, 3, 4, 5, 6], dtype=torch.long)
offsets = torch.tensor([0, 2], dtype=torch.long)
embedding = torch.nn.EmbeddingBag(7, 8, mode="mean")
vectors = embedding(ids, offsets) # [2, 8]
head = torch.nn.Linear(8, 2)
logits = head(vectors) # [2, 2]Check your understanding
The example pools [2,3] and [4,5,6] into two vectors. A third sentence needs one more start offset, not padding to the longest sentence.
A pooled baseline is cheap and inspectable, but “dog bites person” and “person bites dog” have the same bag of words. Use an order-aware contextual model when that distinction matters.
Fine-tuning: compare a new head alone, the head plus selected late blocks, or all layers. Build the optimizer after deciding what is trainable. Dynamic padding reduces unused positions; it does not eliminate every padded computation.
Common mistakes
- Building a custom vocabulary from validation/test text; learn it from training text only.
- Treating a larger token ID as a more important word.
- Averaging embeddings across padding or leaving empty text undefined.
- Assuming class weights replace per-class evaluation; inspect recall and macro F1.
- Saving a model without its vocabulary/tokenizer, label order and length policy.