Conference paper · 2019
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova
11 sections · about 46 minutes · source text with editorial apparatus kept separate