[CME295] 1. Transformer
[CME295] 1. Transformer
1. NLP Overview
1.1 NLP Tasks
- Classification
- Sentiment extraction, Intend detection, Language detection, Topic Modeling
- Multi-Classification
- Part of speech tagging, Named entity recognition, Dependency parsing, Constituency Parsing
- Generation
- Machine translation, Question answering, Summerization, Text generation
1.2 NLP Metrics
- Classification
- Accuracy
- Precision
- Recall
- F1 score
- Generation
- BLEU: 생성 문장에 참조 문장의 단어가 얼마나 많은지 (≒Precision)
- ROUGE: 참조 문장의 단어를 모델이 얼마나 많이 생성했는지 (≒Recall)
- Perplexity(PPL): 모델이 다음 단어를 얼마나 확신하지 못하는지 (확신↑ → PPL↓)
2. Tokenization
- Tokenization: 텍스트를 토큰 단위로 쪼개는 과정
| Method | Pros | Cons |
|---|---|---|
| Word-level | Simple Interpretable | Risk of OOV Does not leverage knowledge of root |
| Subword-level e.g. WordPiece, BPE | Leverages common prefixes and suffixes Learned from the data | Risk of OOV, though less than word-level |
| Character-level | Small chance of OOV RoBUsT tO CASinG anD MIspeliNGs | Makes computations slower Embeddings not interpretable |
3. Word Representation: Word2vec
4. RNNs
5. Self-attention Mechanism
6. Transformer Architecture
This post is licensed under CC BY 4.0 by the author.
