Work overview

Section 07 of 11

Conclusion

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019

Contents

Section 07 of 11

  1. 01Abstract
  2. 02Introduction
  3. 03Related Work
  4. 04BERT
  5. 05Experiments
  6. 06Ablation Studies
  7. 07Conclusion
  8. 08References
  9. 09Additional Details for BERT
  10. 10Detailed Experimental Setup
  11. 11Additional Ablation Studies
Text size
Work overview

Section 7 of 11

Conclusion

Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · about 1 minutes

6 Conclusion

Recent empirical improvements due to transfer learning with language models have demonstrated that rich, unsupervised pre-training is an integral part of many language understanding systems. In particular, these results enable even low-resource tasks to benefit from deep unidirectional architectures. Our major contribution is further generalizing these findings to deep bidirectional architectures, allowing the same pre-trained model to successfully tackle a broad set of NLP tasks.