The official implementation of the ICML 2023 paper OFQ-ViT
-
Updated
Oct 3, 2023 - Python
The official implementation of the ICML 2023 paper OFQ-ViT
From-scratch structured attention-head pruning — reproduces Michel et al. on GPT-2, shows where the method fails on a distilled model, and adds redundancy-scaling and layer-profile analyses.
Papers for deep neural network compression and acceleration
Add a description, image, and links to the model-compression-papers topic page so that developers can more easily learn about it.
To associate your repository with the model-compression-papers topic, visit your repo's landing page and select "manage topics."