Conference paper · 2019

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova

11 sections · about 46 minutes · source text with editorial apparatus kept separate