Pruning using the Lottery Ticket Hypothesis
Testing the Lottery Ticket Hypothesis — the claim that a randomly initialized dense network contains a much smaller subnetwork which, if trained in isolation from the same initial weights, can match the full network's accuracy.
Iterative magnitude pruning
Starting from a dense network with initial weights , each round:
- Trains the current network to convergence.
- Prunes the of remaining weights with the smallest magnitude, producing a binary mask .
- Resets the surviving weights back to their original values from (not their trained values) — this "rewinding" step is what distinguishes a winning ticket from ordinary pruning.
The resulting sparse network at round is
where denotes elementwise multiplication and is the cumulative mask after rounds of pruning.
What "winning" means
A subnetwork is a winning ticket if, trained in isolation, it reaches test accuracy within of the original dense network at the same number of training iterations :
Result
Iterative pruning consistently finds subnetworks at 10–20% of the original parameter count that match dense-network accuracy — but only when rewound to the original initialization; pruning to the same sparsity without rewinding degrades accuracy noticeably, matching the hypothesis's central claim.
Source code: github.com/aarondavis-git/Pruning-LotteryTicketHypothesis