An encoder-only transformer pretrained by masking tokens and predicting them from both sides. Showed that one pretrained model plus a small head beats task-specific architectures.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
specializesis a specific case of Transformerencoder only, trained by masking rather than predicting
Referenced by
alternative-toAutoregressive Language Model is a competing approach to thisgenerates continuations rather than encoding a whole span
References
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding — Devlin, Chang, Lee, Toutanova — 2018 · link