Skip to content

Topic

Depth pruning

Model compression that deletes whole transformer blocks, shortening the network so that each removed block takes one stage of computation out of inference.

Current clusters