When BERT Plays The Lottery, All Tickets Are Winning (Paper Explained)

BERT is a giant model. Turns out you can prune away many of its components and it still works. This paper analyzes BERT pruning in light of the Lottery Ticket Hypothesis and finds that even the “bad” lottery tickets can be fine-tuned to good accuracy.

OUTLINE:
0:00 – Overview
1:20 – BERT
3:20 – Lottery Ticket Hypothesis
13:00 – Paper Abstract
18:00 – Pruning BERT
23:00 – Experiments
50:00 – Conclusion

ML Street Talk Channel:

Abstract:
Much of the recent success in NLP is due to the large Transformer-based models such as BERT (Devlin et al, 2019). However, these models have been shown to be reducible to a smaller number of self-attention heads and layers. We consider this phenomenon from the perspective of the lottery ticket hypothesis. For fine-tuned BERT, we show that (a) it is possible to find a subnetwork of elements that achieves performance comparable with that of the full model, and (b) similarly-sized subnetworks sampled from the rest of the model perform worse. However, the “bad” subnetworks can be fine-tuned separately to achieve only slightly worse performance than the “good” ones, indicating that most weights in the pre-trained BERT are potentially useful. We also show that the “good” subnetworks vary considerably across GLUE tasks, opening up the possibilities to learn what knowledge BERT actually uses at inference time.

Authors: Sai Prasanna, Anna Rogers, Anna Rumshisky

Links:
YouTube:
Twitter:
BitChute:
Minds:

source

What's Hot

Australia’s biggest bank cut staff for AI, then it backtracked – and it’s one of many scrapping plans for automated customer support teams

The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward – Takara TLDR

Bulgarian Doctoral Student Anna-Maria Halacheva Recognized by European Commission – Novinite.com

When BERT Plays the Lottery, All Tickets Are Winning (Paper Explained)

AGI is not coming!

Context Rot: How Increasing Input Tokens Impacts LLM Performance (Paper Analysis)

Energy-Based Transformers are Scalable Learners and Thinkers (Paper Review)

Long-Lost Painting By Rubens From 1613 Discovered in Paris Mansion

Ken Griffin Loves Pollock’s Blue Poles So Much He Tried to Buy it

Sally Mann Says Her Black Men Photos Are ‘Problematic’ in Hindsight

NeueHouse, a Hot Spot for Art Events, Files for Bankruptcy

Australia’s biggest bank cut staff for AI, then it backtracked – and it’s one of many scrapping plans for automated customer support teams

The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward – Takara TLDR

Bulgarian Doctoral Student Anna-Maria Halacheva Recognized by European Commission – Novinite.com

What's Hot

When BERT Plays the Lottery, All Tickets Are Winning (Paper Explained)

Related Posts

Subscribe to Updates