COMP9019046 Semantic Representation, Tokenization
SB
Published · 47 slides · 0 views
1 / 1
Description
COMP9019046 Semantic Representation, Tokenization Embeddings Session 2 Part 1 Text Preprocessing Tokenization, normalization, stopword removal, stemming, lemmatization, regex, edit distance, dan preprocessing pipeline Text Preprocessing:
Related Topics
Share
Embed code
Download this presentation From Below
"COMP9019046 Semantic Representation, Tokenization" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.
Presentation Transcript
01
COMP9019046
Semantic Representation, Tokenization & Embeddings
Session 2<br>
Semantic Representation, Tokenization & Embeddings
Session 2<br>
02
Part 1
Text Preprocessing
Tokenization, normalization, stopword removal, stemming,
lemmatization, regex, edit distance, dan preprocessing pipeline<br>
Text Preprocessing
Tokenization, normalization, stopword removal, stemming,
lemmatization, regex, edit distance, dan preprocessing pipeline<br>
03
Text Preprocessing: What Are We Trying to Preserve? Principle
Preprocessing is task-dependent: preserve information that matters to the downstream task. Knowledge & Information Retrieval • Session 02<br>
Preprocessing is task-dependent: preserve information that matters to the downstream task. Knowledge & Information Retrieval • Session 02<br>
04
Tokenization — Word-Level Example Raw text I love NLP! Word tokens [I] [love] [NLP] [!] Words are intuitive tokens for many classical NLP pipelines.
Punctuation can be treated as its own token.
Whitespace is a simple signal, but languages differ.
Social media adds hashtags, mentions, emojis, URLs, and repeated characters. Knowledge & Information Retrieval • Session 02<br>
Punctuation can be treated as its own token.
Whitespace is a simple signal, but languages differ.
Social media adds hashtags, mentions, emojis, URLs, and repeated characters. Knowledge & Information Retrieval • Session 02<br>
05
Subword Tokenization — Why Do We Need It? Subword methods help reduce the out-of-vocabulary problem.
A rare word can be represented using pieces seen during training.
Exact pieces depend on the tokenizer/vocabulary; examples are illustrative.
Subword tokenization is common in transformer-based NLP systems. Knowledge & Information Retrieval • Session 02<br>
A rare word can be represented using pieces seen during training.
Exact pieces depend on the tokenizer/vocabulary; examples are illustrative.
Subword tokenization is common in transformer-based NLP systems. Knowledge & Information Retrieval • Session 02<br>
06
Tokenization: A Small Hands-On Task Task: tokenize this sentence “I’m learning NLP at BINUS University.” Knowledge & Information Retrieval • Session 02<br>
07
Normalization — Lowercasing Benefit: reduces vocabulary sparsity.
Risk: case can carry information.
“Apple” can refer to a company while “apple” can refer to a fruit.
Named entities, acronyms, product names, and sentence starts may be affected. Knowledge & Information Retrieval • Session 02<br>
Risk: case can carry information.
“Apple” can refer to a company while “apple” can refer to a fruit.
Named entities, acronyms, product names, and sentence starts may be affected. Knowledge & Information Retrieval • Session 02<br>
08
Normalization — Punctuation, Numbers & URLs Replacing patterns with placeholders can preserve useful information while reducing sparsity.
Removing every number may be harmful for financial, scientific, medical, or date-related tasks.
Removing URLs can be useful for topic classification but harmful if the URL itself carries information. Knowledge & Information Retrieval • Session 02<br>
Removing every number may be harmful for financial, scientific, medical, or date-related tasks.
Removing URLs can be useful for topic classification but harmful if the URL itself carries information. Knowledge & Information Retrieval • Session 02<br>
09
Emojis and Social-Media Text<br>
10
Stopword Removal — Example Original the cat is on the table After removal cat table Stopwords are high-frequency words that may contribute little to some tasks.
Removing them can reduce vocabulary size.
But “not” is crucial in sentiment: “good” ≠ “not good”.
Function words can matter in syntax, question answering, and information extraction. Knowledge & Information Retrieval • Session 02<br>
Removing them can reduce vocabulary size.
But “not” is crucial in sentiment: “good” ≠ “not good”.
Function words can matter in syntax, question answering, and information extraction. Knowledge & Information Retrieval • Session 02<br>
11
Stopwords — Task Dependency Same preprocessing ≠ best preprocessing for every NLP task Knowledge & Information Retrieval • Session 02<br>
12
Stemming — Fast Morphological Reduction Stemming usually applies rules to reduce word forms.
It can be fast and useful for retrieval or rough normalization.
Output may not be a valid dictionary word.
Different stemmers can produce different outputs. Use when
Speed and rough normalization are more important than linguistically valid word forms. Knowledge & Information Retrieval • Session 02<br>
It can be fast and useful for retrieval or rough normalization.
Output may not be a valid dictionary word.
Different stemmers can produce different outputs. Use when
Speed and rough normalization are more important than linguistically valid word forms. Knowledge & Information Retrieval • Session 02<br>
13
Lemmatization — Linguistically Valid Base Forms Lemmatization aims to return a dictionary/base form.
It uses vocabulary and grammatical information.
It is more linguistically informed than simple truncation.
Context and part-of-speech information can affect the lemma. Compare
Stemming asks “what string remains?” Lemmatization asks “what valid base form does this word represent?” Knowledge & Information Retrieval • Session 02<br>
It uses vocabulary and grammatical information.
It is more linguistically informed than simple truncation.
Context and part-of-speech information can affect the lemma. Compare
Stemming asks “what string remains?” Lemmatization asks “what valid base form does this word represent?” Knowledge & Information Retrieval • Session 02<br>
14
Stemming vs Lemmatization — When to Choose? Knowledge & Information Retrieval • Session 02 Knowledge & Information Retrieval • Session 02<br>
15
Regex — Pattern Matching in Text \d+ → one or more digits \w+ → one or more word characters #\w+ → hashtag-like pattern Knowledge & Information Retrieval • Session 02<br>
16
Regex — Extract Instead of Delete Input Email me at hello@binus.ac.id Pattern [\w.+-]+@[\w.-]+\.\w+ Output hello@binus.ac.id Regex can extract information, not only remove it.
Useful targets: emails, URLs, hashtags, mentions, dates, IDs, repeated punctuation.
Always validate patterns on representative data. Knowledge & Information Retrieval • Session 02<br>
Useful targets: emails, URLs, hashtags, mentions, dates, IDs, repeated punctuation.
Always validate patterns on representative data. Knowledge & Information Retrieval • Session 02<br>
17
Regex Cleaning Pipeline Raw text Detect URLs Detect mentions Detect hashtags Normalize spaces Clean text “Hi!!! Visit https://x.com @student #NLP” “Hi Visit [URL] [MENTION] [HASHTAG]” Whether to replace patterns or remove them depends on the task.
Placeholders can retain information type without retaining the exact string. Knowledge & Information Retrieval • Session 02<br>
Placeholders can retain information type without retaining the exact string. Knowledge & Information Retrieval • Session 02<br>
18
Edit Distance — Intuition kitten sitten sittin sitting kitten → sitten : substitution = 1 sitten → sittin : substitution = 1 sittin → sitting : insertion = 1 Total Levenshtein distance = 3 Allowed operations: insertion, deletion, substitution.
Lower distance means fewer edits are needed.
Useful for noisy text, typo correction, deduplication, and record linkage. Knowledge & Information Retrieval • Session 02 18 18<br>
Lower distance means fewer edits are needed.
Useful for noisy text, typo correction, deduplication, and record linkage. Knowledge & Information Retrieval • Session 02 18 18<br>
19
Edit Distance — Spell Correction Example Edit distance alone can produce multiple candidates with the same distance.
“wll” → “will” and “wll” → “wall” may both be close.
Context or a language model can decide which candidate is more plausible. Knowledge & Information Retrieval • Session 02<br>
“wll” → “will” and “wll” → “wall” may both be close.
Context or a language model can decide which candidate is more plausible. Knowledge & Information Retrieval • Session 02<br>
20
Integrated Text Preprocessing Pipeline — Practice Raw text Normalize Tokenize Stopwords Lemma/stem Regex Final tokens “OMG!!! I’m running to the stores 😍 #shopping” Knowledge & Information Retrieval • Session 02 21<br>
21
Part 2
Semantic Representation
Semantic representation, BoW, TF-IDF, contoh perhitungan TF-IDF,
document vector, cosine similarity, dan semantic gap<br>
Semantic Representation
Semantic representation, BoW, TF-IDF, contoh perhitungan TF-IDF,
document vector, cosine similarity, dan semantic gap<br>
22
Why Representation Matters Core idea
A good representation makes the information relevant to the task easier for the model to learn or compare. Knowledge & Information Retrieval • Session 02<br>
A good representation makes the information relevant to the task easier for the model to learn or compare. Knowledge & Information Retrieval • Session 02<br>
23
Tokenization: The First Representation Step Text “I love semantic search!” Tokens [I] [love] [semantic] [search] [!] Tokenization splits text into smaller units called tokens.
A token is not necessarily a complete word.
Common choices: word-level, subword, and character-level tokenization.
Tokenization is foundational because later representations operate on these units. Knowledge & Information Retrieval • Session 02<br>
A token is not necessarily a complete word.
Common choices: word-level, subword, and character-level tokenization.
Tokenization is foundational because later representations operate on these units. Knowledge & Information Retrieval • Session 02<br>
24
1. What Is Semantic Representation? Human language “I need a cheap phone” Numerical representation [0.21, −0.08, 0.73, …] Semantic representation = a numerical form intended to preserve useful information about meaning, usage, or relationships.
Computers can then compare, classify, cluster, rank, or retrieve text.
Representation can be sparse (BoW/TF-IDF) or dense (embeddings).
Different tasks require different representations. Knowledge & Information Retrieval • Session 02 4<br>
Computers can then compare, classify, cluster, rank, or retrieve text.
Representation can be sparse (BoW/TF-IDF) or dense (embeddings).
Different tasks require different representations. Knowledge & Information Retrieval • Session 02 4<br>
25
2. Why Representation Matters Core idea
A good representation makes the information relevant to the task easier for the model to learn or compare. Knowledge & Information Retrieval • Session 02 5<br>
A good representation makes the information relevant to the task easier for the model to learn or compare. Knowledge & Information Retrieval • Session 02 5<br>
26
3. Tokenization: The First Representation Step Text “I love semantic search!” Tokens [I] [love] [semantic] [search] [!] Tokenization splits text into smaller units called tokens.
A token is not necessarily a complete word.
Common choices: word-level, subword, and character-level tokenization.
Tokenization is foundational because later representations operate on these units. Knowledge & Information Retrieval • Session 02 6<br>
A token is not necessarily a complete word.
Common choices: word-level, subword, and character-level tokenization.
Tokenization is foundational because later representations operate on these units. Knowledge & Information Retrieval • Session 02 6<br>
27
4. Three Tokenization Strategies Example: “I’m loving NLP!” may become several subword tokens depending on the tokenizer. Therefore token counts can differ substantially from word counts. 7<br>
28
5. Token → Token ID → Vector ID 1842 → embedding matrix row 1842 → [0.21, −0.08, 0.73, …] Token ID is categorical: 3911 is not “more semantic” than 572.
The embedding vector is where learned numerical relationships can be represented.
Think of token IDs as addresses; embeddings are the content stored at those addresses. Knowledge & Information Retrieval • Session 02 8<br>
The embedding vector is where learned numerical relationships can be represented.
Think of token IDs as addresses; embeddings are the content stored at those addresses. Knowledge & Information Retrieval • Session 02 8<br>
29
6. Representation Spectrum: From Counts to Meaning BoW TF-IDF Word Embeddings Sentence Embeddings Document Embeddings 9<br>
30
7. BoW: Simple but Sparse BoW ignores grammar and word order.
“car” and “automobile” are different dimensions.
High-dimensional sparse vectors can still work well for small or classical IR tasks. Good for: baseline search, interpretable term matching Weak for: semantic matching and synonyms Knowledge & Information Retrieval • Session 02 10<br>
“car” and “automobile” are different dimensions.
High-dimensional sparse vectors can still work well for small or classical IR tasks. Good for: baseline search, interpretable term matching Weak for: semantic matching and synonyms Knowledge & Information Retrieval • Session 02 10<br>
31
8. TF-IDF: Weight Important Terms TF(t,d) = count(t in d) / total terms in d IDF(t) = log(N / DF(t)) TF-IDF(t,d) = TF(t,d) × IDF(t) Common words appearing in many documents receive low IDF.
Rare, informative terms receive higher weight.
Still a lexical representation: it does not inherently know that “car” and “automobile” are synonyms. 11<br>
Rare, informative terms receive higher weight.
Still a lexical representation: it does not inherently know that “car” and “automobile” are synonyms. 11<br>
32
9. The Semantic Gap Lexical matching asks: “Do the same words occur?” Semantic matching asks: “Do these texts mean something similar?” 12<br>
33
Part 3
Embeddings
Word embeddings, Word2Vec, CBOW vs Skip-gram, embedding space, sentence embeddings, document embeddings, static vs contextual embeddings, dan Transformer representation.<br>
Embeddings
Word embeddings, Word2Vec, CBOW vs Skip-gram, embedding space, sentence embeddings, document embeddings, static vs contextual embeddings, dan Transformer representation.<br>
34
Word Embeddings word → dense vector coffee → [0.21, −0.08, 0.73, 0.44, …] A word embedding represents a word as a dense numerical vector.
Dimensions are learned features; they are not normally human-readable labels such as “is-a-food”.
Words occurring in similar contexts tend to develop similar representations.
Embeddings can support similarity, clustering, classification, recommendation, and retrieval. Knowledge & Information Retrieval • Session 02 13<br>
Dimensions are learned features; they are not normally human-readable labels such as “is-a-food”.
Words occurring in similar contexts tend to develop similar representations.
Embeddings can support similarity, clustering, classification, recommendation, and retrieval. Knowledge & Information Retrieval • Session 02 13<br>
35
Word2Vec: Learning from Context CBOW context words → target word Skip-gram target word → context words CBOW predicts a target word from nearby context words.
Skip-gram predicts nearby context words from a target word.
The learned hidden representation becomes a useful word vector.
Example: “I drink hot coffee every morning.” Context can help learn relationships among drink, coffee, cup, morning, etc. 14<br>
Skip-gram predicts nearby context words from a target word.
The learned hidden representation becomes a useful word vector.
Example: “I drink hot coffee every morning.” Context can help learn relationships among drink, coffee, cup, morning, etc. 14<br>
36
Word Embedding Space — Conceptual Example coffee tea cake car bus doctor 2D picture is only a visualization; real embeddings may have hundreds of dimensions. Nearby vectors can indicate related usage.
Distance is not automatically “meaning” unless the model and training objective support it.
Different embedding models can create different spaces. 15<br>
Distance is not automatically “meaning” unless the model and training objective support it.
Different embedding models can create different spaces. 15<br>
37
Cosine Similarity: Compare Directions cos(A,B) = (A · B) / (||A|| ||B||) A=[1,2], B=[2,3] A·B = 1×2 + 2×3 = 8 ||A||=√5, ||B||=√13 cos(A,B)=8/√65≈0.992 Interpretation
A value close to 1 means the vectors point in similar directions. In many embedding retrieval systems, higher cosine similarity means stronger semantic similarity. 16<br>
A value close to 1 means the vectors point in similar directions. In many embedding retrieval systems, higher cosine similarity means stronger semantic similarity. 16<br>
38
Cosine Similarity — Calculate Every Component cos(A,B) = (A·B) / (||A|| ||B||) A = [1, 1, 1, 1, 0, 1, 1, 2] B = [1, 0, 0, 1, 1, 0, 1, 0] A·B = 1+0+0+1+0+0+1+0 = 3 ||A|| = √(1+1+1+1+0+1+1+4) = √10 ||B|| = √(1+0+0+1+1+0+1+0) = √4 = 2 cos(A,B) = 3 / (√10 × 2) ≈ 0.4743 Interpretation
The result 0.4743 means the vectors share some directional similarity, but they are far from identical. The score is a similarity measure—not a probability. 17<br>
The result 0.4743 means the vectors share some directional similarity, but they are far from identical. The score is a similarity measure—not a probability. 17<br>
39
Cosine Similarity — Use It to Rank Search Results Higher cosine → higher rank (when using cosine as the ranking score) Step 1: represent query and documents in the same vector space.
Step 2: calculate cosine similarity for each candidate.
Step 3: sort from highest to lowest.
Step 4: return top-k results.
Question: what happens when the query and relevant document use synonyms rather than the same terms? This motivates semantic embeddings. 18<br>
Step 2: calculate cosine similarity for each candidate.
Step 3: sort from highest to lowest.
Step 4: return top-k results.
Question: what happens when the query and relevant document use synonyms rather than the same terms? This motivates semantic embeddings. 18<br>
40
Sentence Embeddings Sentence Token representations Pooling / encoder One vector 19<br>
41
How Can We Build a Sentence Embedding? 20<br>
42
Document Embeddings Document “A 12-page report about electric vehicles…” Document vector [0.11, 0.42, …] Long documents may need chunking before embedding.
A document embedding should preserve information relevant to the intended task.
Examples: document clustering, semantic search, recommendation, duplicate detection.
For RAG, embedding chunks is often more practical than embedding an entire long document as one vector. 21<br>
A document embedding should preserve information relevant to the intended task.
Examples: document clustering, semantic search, recommendation, duplicate detection.
For RAG, embedding chunks is often more practical than embedding an entire long document as one vector. 21<br>
43
Word vs Sentence vs Document Embeddings Granularity matters: a good word embedding does not automatically become a good document embedding.
Longer text introduces aggregation, context, chunking, and information-loss issues. 22<br>
Longer text introduces aggregation, context, chunking, and information-loss issues. 22<br>
44
Static vs Contextual Embeddings Static: meaning is relatively fixed per word Contextual: representation changes with surrounding text 23<br>
45
1. What is the difference between a token ID and an embedding vector?
2. Why can “car” and “automobile” be difficult for BoW/TF-IDF but easier for embeddings?
3. When would you prefer a sentence embedding over a word embedding?
4. Why is chunking important for document retrieval?
5. What does a high cosine similarity score tell you—and what does it NOT tell you? Practice and discussion<br>
2. Why can “car” and “automobile” be difficult for BoW/TF-IDF but easier for embeddings?
3. When would you prefer a sentence embedding over a word embedding?
4. Why is chunking important for document retrieval?
5. What does a high cosine similarity score tell you—and what does it NOT tell you? Practice and discussion<br>
46
Paper Review Exercise
A Novel Text Mining Approach for Mental Health Prediction Using Bi-LSTM and BERT Model - PMC What is the research problem?
What datasets were used?
How was the social-media data collected and labeled?
What text representation techniques were compared?
Why did the authors use BERT?
What is the role of Bi-LSTM and knowledge distillation?
Is 98% accuracy sufficient evidence that the proposed model is effective?
What is one major limitation of this study and how would you improve it? Work individually. Read the Abstract, Introduction, Methodology, Experimental Results, and Conclusion sections, then answer the questions below.
Your review should focus not only on what the authors did, but also on why they did it and whether the methodology and results are convincing.<br>
A Novel Text Mining Approach for Mental Health Prediction Using Bi-LSTM and BERT Model - PMC What is the research problem?
What datasets were used?
How was the social-media data collected and labeled?
What text representation techniques were compared?
Why did the authors use BERT?
What is the role of Bi-LSTM and knowledge distillation?
Is 98% accuracy sufficient evidence that the proposed model is effective?
What is one major limitation of this study and how would you improve it? Work individually. Read the Abstract, Introduction, Methodology, Experimental Results, and Conclusion sections, then answer the questions below.
Your review should focus not only on what the authors did, but also on why they did it and whether the methodology and results are convincing.<br>
47
Thank You<br>