Post

[CME295] 1. Transformer

[CME295] 1. Transformer

1. NLP Overview

1.1 NLP Tasks

NLP Tasks

  • Classification
    • Sentiment extraction, Intend detection, Language detection, Topic Modeling
  • Multi-Classification
    • Part of speech tagging, Named entity recognition, Dependency parsing, Constituency Parsing
  • Generation
    • Machine translation, Question answering, Summerization, Text generation

1.2 NLP Metrics

  • Classification
    • Accuracy
    • Precision
    • Recall
    • F1 score
  • Generation
    • BLEU: 생성 문장에 참조 문장의 단어가 얼마나 많은지 (≒Precision)
    • ROUGE: 참조 문장의 단어를 모델이 얼마나 많이 생성했는지 (≒Recall)
    • Perplexity(PPL): 모델이 다음 단어를 얼마나 확신하지 못하는지 (확신↑ → PPL↓)

2. Tokenization

  • Tokenization: 텍스트를 토큰 단위로 쪼개는 과정
MethodProsCons
Word-levelSimple
Interpretable
Risk of OOV
Does not leverage knowledge of root
Subword-level
e.g. WordPiece, BPE
Leverages common prefixes and suffixes
Learned from the data
Risk of OOV, though less than word-level
Character-levelSmall chance of OOV
RoBUsT tO CASinG anD MIspeliNGs
Makes computations slower
Embeddings not interpretable

3. Word Representation: Word2vec

4. RNNs

5. Self-attention Mechanism

6. Transformer Architecture

This post is licensed under CC BY 4.0 by the author.