[Doc] Remove legacy doc (#4623)

88ffe908 · J-shang · GitHub · 006e1a01 · 006e1a01 · 88ffe908
Unverified Commit 88ffe908 authored Mar 09, 2022 by J-shang Committed by GitHub Mar 09, 2022
19 changed files
--- a/docs/source/compression.rst
+++ b/docs/source/compression.rst
-#################
-Model Compression
-#################
-Deep neural networks (DNNs) have achieved great success in many tasks. However, typical neural networks are both
-computationally expensive and energy intensive, can be difficult to be deployed on devices with low computation
-resources or with strict latency requirements. Therefore, a natural thought is to perform model compression to
-reduce model size and accelerate model training/inference without losing performance significantly. Model compression
-techniques can be divided into two categories: pruning and quantization. The pruning methods explore the redundancy
-in the model weights and try to remove/prune the redundant and uncritical weights. Quantization refers to compressing
-models by reducing the number of bits required to represent weights or activations.
-NNI provides an easy-to-use toolkit to help user design and use model pruning and quantization algorithms.
-It supports Tensorflow and PyTorch with unified interface.
-For users to compress their models, they only need to add several lines in their code.
-There are some popular model compression algorithms built-in in NNI.
-Users could further use NNI's auto tuning power to find the best compressed model,
-which is detailed in Auto Model Compression.
-On the other hand, users could easily customize their new compression algorithms using NNI's interface.
-For details, please refer to the following tutorials:
-..  toctree::
-    :maxdepth: 2
-    Overview <compression/overview>
-    Pruning <compression/pruning>
-    Quantization <compression/quantization>
-    Advanced Usage <compression/advanced_usage>
-    Reference <compression/reference>
--- a/docs/source/compression/advanced_usage.rst
+++ b/docs/source/compression/advanced_usage.rst
@@ -6,6 +6,3 @@ Advanced Usage
    Customize Scheduled Pruning Process <pruning_scheduler>
    Utilities <compression_utils>
-    Framework (legacy) <legacy_framework>
-    Customize a New Algorithm (legacy) <legacy_customize_compressor>
-    Automatic Model Compression (legacy) <legacy_autocompression>
--- a/docs/source/reference/compression_config_list.rst
+++ b/docs/source/reference/compression_config_list.rst
--- a/docs/source/compression/overview.rst
+++ b/docs/source/compression/overview.rst
-Model Compression Overview
+Model Compression with NNI
 ==========================
-Deep neural networks (DNNs) have achieved great success in many tasks like embedded development and scenarios that needs rapid feedbacks.
+.. toctree::
+    :hidden:
+    :maxdepth: 2
+    Pruning <pruning>
+    Quantization <quantization>
+    Config Specification <compression_config_list>
+    Advanced Usage <advanced_usage>
+.. attention::
+  NNI's model pruning framework has been upgraded to a more powerful version (named pruning v2 before nni v2.6).
+  The old version (`named pruning before nni v2.6 <https://nni.readthedocs.io/en/v2.6/Compression/pruning.html>`_) will be out of maintenance. If for some reason you have to use the old pruning,
+  v2.6 is the last nni version to support old pruning version.
+.. Using rubric to prevent the section heading to be include into toc
+.. rubric:: Overview
+Deep neural networks (DNNs) have achieved great success in many tasks like computer vision, nature launguage processing, speech processing.
 However, typical neural networks are both computationally expensive and energy-intensive,
 which can be difficult to be deployed on devices with low computation resources or with strict latency requirements.
 Therefore, a natural thought is to perform model compression to reduce model size and accelerate model training/inference without losing performance significantly.
 Model compression techniques can be divided into two categories: pruning and quantization.
 The pruning methods explore the redundancy in the model weights and try to remove/prune the redundant and uncritical weights.
-Quantization refers to compressing models by reducing the number of bits required to represent weights or activations functions.
+Quantization refers to compress models by reducing the number of bits required to represent weights or activations.
 We further elaborate on the two methods, pruning and quantization, in the following chapters. Besides, the figure below visualizes the difference between these two methods.
 .. image:: ../../img/prune_quant.jpg
   :target: ../../img/prune_quant.jpg
   :scale: 40%
   :alt:
 NNI provides an easy-to-use toolkit to help users design and use model pruning and quantization algorithms.
 For users to compress their models, they only need to add several lines in their code.
 There are some popular model compression algorithms built-in in NNI.
@@ -33,8 +49,7 @@ There are several core features supported by NNI model compression:
 * Concise interface for users to customize their own compression algorithms.
-Compression Pipeline
+.. rubric:: Compression Pipeline
--------------------
 .. image:: ../../img/compression_flow.jpg
   :target: ../../img/compression_flow.jpg
@@ -49,13 +64,7 @@ If users want to apply both, a sequential mode is recommended as common practise
  The interface and APIs are unified for both PyTorch and TensorFlow. Currently only PyTorch version has been supported, and TensorFlow version will be supported in future.
-Supported Algorithms
+.. rubric:: Supported Pruning Algorithms
--------------------
-The supported model compression algorithms include pruning algorithms and quantization algorithms.
-Pruning Algorithms
-^^^^^^^^^^^^^^^^^^
 Pruning algorithms compress the original network by removing redundant weights or channels of layers, which can reduce model complexity and mitigate the over-fitting issue.
@@ -99,8 +108,7 @@ Pruning algorithms compress the original network by removing redundant weights o
     - Movement Pruning: Adaptive Sparsity by Fine-Tuning `Reference Paper <https://arxiv.org/abs/2005.07683>`__
-Quantization Algorithms
+.. rubric:: Supported Quantization Algorithms
-^^^^^^^^^^^^^^^^^^^^^^^
 Quantization algorithms compress the original network by reducing the number of bits required to represent weights or activations, which can reduce the computations and the inference time.
@@ -124,21 +132,19 @@ Quantization algorithms compress the original network by reducing the number of
     - Post training quantizaiton. Collect quantization information during calibration with observers.
-Model Speedup
+.. rubric:: Model Speedup
-------------
 The final goal of model compression is to reduce inference latency and model size.
 However, existing model compression algorithms mainly use simulation to check the performance (e.g., accuracy) of compressed model.
 For example, using masks for pruning algorithms, and storing quantized values still in float32 for quantization algorithms.
-Given the output masks and quantization bits produced by those algorithms, NNI can really speed up the model. The following figure shows how NNI prunes and speeds up your models. 
+Given the output masks and quantization bits produced by those algorithms, NNI can really speed up the model.
+The following figure shows how NNI prunes and speeds up your models. 
 .. image:: ../../img/pipeline_compress.jpg
   :target: ../../img/pipeline_compress.jpg
   :scale: 40%
   :alt:
 The detailed tutorial of Speed Up Model with Mask can be found :doc:`here <../tutorials/pruning_speed_up>`.
 The detailed tutorial of Speed Up Model with Calibration Config can be found :doc:`here <../tutorials/quantization_speed_up>`.
\ No newline at end of file
--- a/docs/source/compression_zh.rst
+++ b/docs/source/compression_zh.rst
-.. d371fe9f337e7c445c2f3016fc939aaf
+.. 1d14b9d13cdd660f8e9dcb2abed0b185
-#################
 模型压缩
-#################
+========
+..  toctree::
+    :hidden:
+    :maxdepth: 2
+    模型剪枝 <pruning>
+    模型量化 <quantization>
+    用户配置 <compression_config_list>
+    高级用法 <advanced_usage>
 深度神经网络（DNNs）在许多领域都取得了巨大的成功。 然而，典型的神经网络是
 计算和能源密集型的，很难将其部署在计算资源匮乏
@@ -19,14 +27,3 @@ NNI 中也内置了一些主流的模型压缩算法。
 用户可以进一步利用 NNI 的自动调优功能找到最佳的压缩模型，
 该功能在自动模型压缩部分有详细介绍。
 另一方面，用户可以使用 NNI 的接口自定义新的压缩算法。
-详细信息，参考以下教程：
-..  toctree::
-    :maxdepth: 2
-    概述 <compression/overview>
-    模型剪枝 <compression/pruning>
-    模型量化 <compression/quantization>
-    高级用法 <compression/advanced_usage>
-    参考 <compression/reference>
--- a/docs/source/compression/legacy_autocompression.rst
+++ b/docs/source/compression/legacy_autocompression.rst
-Auto Compression with NNI Experiment
-====================================
-If you want to compress your model, but don't know what compression algorithm to choose, or don't know what sparsity is suitable for your model, or just want to try more possibilities, auto compression may help you.
-Users can choose different compression algorithms and define the algorithms' search space, then auto compression will launch an NNI experiment and try different compression algorithms with varying sparsity automatically. 
-Of course, in addition to the sparsity rate, users can also introduce other related parameters into the search space.
-If you don't know what is search space or how to write search space, `this <./Tutorial/SearchSpaceSpec.rst>`__ is for your reference.
-Auto compression using experience is similar to the NNI experiment in python.
-The main differences are as follows:
-* Use a generator to help generate search space object.
-* Need to provide the model to be compressed, and the model should have already been pre-trained.
-* No need to set ``trial_command``, additional need to set ``auto_compress_module`` as ``AutoCompressionExperiment`` input.
-.. note::
-    Auto compression only supports TPE Tuner, Random Search Tuner, Anneal Tuner, Evolution Tuner right now.
-Generate search space
---------------------
-Due to the extensive use of nested search space, we recommend a using generator to configure search space.
-The following is an example. Using ``add_config()`` add subconfig, then ``dumps()`` search space dict.
-.. code-block:: python
-    from nni.algorithms.compression.pytorch.auto_compress import AutoCompressionSearchSpaceGenerator
-    generator = AutoCompressionSearchSpaceGenerator()
-    generator.add_config('level', [
-        {
-            "sparsity": {
-                "_type": "uniform",
-                "_value": [0.01, 0.99]
-            },
-            'op_types': ['default']
-        }
-    ])
-    generator.add_config('qat', [
-    {
-        'quant_types': ['weight', 'output'],
-        'quant_bits': {
-            'weight': 8,
-            'output': 8
-        },
-        'op_types': ['Conv2d', 'Linear']
-    }])
-    search_space = generator.dumps()
-Now we support the following pruners and quantizers:
-.. code-block:: python
-    PRUNER_DICT = {
-        'level': LevelPruner,
-        'slim': SlimPruner,
-        'l1': L1FilterPruner,
-        'l2': L2FilterPruner,
-        'fpgm': FPGMPruner,
-        'taylorfo': TaylorFOWeightFilterPruner,
-        'apoz': ActivationAPoZRankFilterPruner,
-        'mean_activation': ActivationMeanRankFilterPruner
-    }
-    QUANTIZER_DICT = {
-        'naive': NaiveQuantizer,
-        'qat': QAT_Quantizer,
-        'dorefa': DoReFaQuantizer,
-        'bnn': BNNQuantizer
-    }
-Provide user model for compression
----------------------------------
-Users need to inherit ``AbstractAutoCompressionModule`` and override the abstract class function.
-.. code-block:: python
-    from nni.algorithms.compression.pytorch.auto_compress import AbstractAutoCompressionModule
-    class AutoCompressionModule(AbstractAutoCompressionModule):
-        @classmethod
-        def model(cls) -> nn.Module:
-            ...
-            return _model
-        @classmethod
-        def evaluator(cls) -> Callable[[nn.Module], float]:
-            ...
-            return _evaluator
-Users need to implement at least ``model()`` and ``evaluator()``.
-If you use iterative pruner, you need to additional implement ``optimizer_factory()``, ``criterion()`` and ``sparsifying_trainer()``.
-If you want to finetune the model after compression, you need to implement ``optimizer_factory()``, ``criterion()``, ``post_compress_finetuning_trainer()`` and ``post_compress_finetuning_epochs()``.
-The ``optimizer_factory()`` should return a factory function, the input is an iterable variable, i.e. your ``model.parameters()``, and the output is an optimizer instance.
-The two kinds of ``trainer()`` should return a trainer with input ``model, optimizer, criterion, current_epoch``.
-The full abstract interface refers to :githublink:`interface.py <nni/algorithms/compression/pytorch/auto_compress/interface.py>`.
-An example of ``AutoCompressionModule`` implementation refers to :githublink:`auto_compress_module.py <examples/model_compress/auto_compress/torch/auto_compress_module.py>`.
-Launch NNI experiment
---------------------
-Similar to launch from python, the difference is no need to set ``trial_command`` and put the user-provided ``AutoCompressionModule`` as ``AutoCompressionExperiment`` input.
-.. code-block:: python
-    from pathlib import Path
-    from nni.algorithms.compression.pytorch.auto_compress import AutoCompressionExperiment
-    from auto_compress_module import AutoCompressionModule
-    experiment = AutoCompressionExperiment(AutoCompressionModule, 'local')
-    experiment.config.experiment_name = 'auto compression torch example'
-    experiment.config.trial_concurrency = 1
-    experiment.config.max_trial_number = 10
-    experiment.config.search_space = search_space
-    experiment.config.trial_code_directory = Path(__file__).parent
-    experiment.config.tuner.name = 'TPE'
-    experiment.config.tuner.class_args['optimize_mode'] = 'maximize'
-    experiment.config.training_service.use_active_gpu = True
-    experiment.run(8088)
--- a/docs/source/compression/legacy_customize_compressor.rst
+++ b/docs/source/compression/legacy_customize_compressor.rst
-Customize New Compression Algorithm
-===================================
-.. contents::
-In order to simplify the process of writing new compression algorithms, we have designed simple and flexible programming interface, which covers pruning and quantization. Below, we first demonstrate how to customize a new pruning algorithm and then demonstrate how to customize a new quantization algorithm.
-**Important Note** To better understand how to customize new pruning/quantization algorithms, users should first understand the framework that supports various pruning algorithms in NNI. Reference :doc:`Framework overview of model compression <legacy_framework>`
-Customize a new pruning algorithm
---------------------------------
-Implementing a new pruning algorithm requires implementing a ``weight masker`` class which shoud be a subclass of ``WeightMasker``\ , and a ``pruner`` class, which should be a subclass ``Pruner``.
-An implementation of ``weight masker`` may look like this:
-.. code-block:: python
-   class MyMasker(WeightMasker):
-       def __init__(self, model, pruner):
-           super().__init__(model, pruner)
-           # You can do some initialization here, such as collecting some statistics data
-           # if it is necessary for your algorithms to calculate the masks.
-       def calc_mask(self, sparsity, wrapper, wrapper_idx=None):
-           # calculate the masks based on the wrapper.weight, and sparsity, 
-           # and anything else
-           # mask = ...
-           return {'weight_mask': mask}
-You can reference nni provided :githublink:`weight masker <nni/algorithms/compression/pytorch/pruning/structured_pruning_masker.py>` implementations to implement your own weight masker.
-A basic ``pruner`` looks likes this:
-.. code-block:: python
-   class MyPruner(Pruner):
-       def __init__(self, model, config_list, optimizer):
-           super().__init__(model, config_list, optimizer)
-           self.set_wrappers_attribute("if_calculated", False)
-           # construct a weight masker instance
-           self.masker = MyMasker(model, self)
-       def calc_mask(self, wrapper, wrapper_idx=None):
-           sparsity = wrapper.config['sparsity']
-           if wrapper.if_calculated:
-               # Already pruned, do not prune again as a one-shot pruner
-               return None
-           else:
-               # call your masker to actually calcuate the mask for this layer
-               masks = self.masker.calc_mask(sparsity=sparsity, wrapper=wrapper, wrapper_idx=wrapper_idx)
-               wrapper.if_calculated = True
-               return masks
-Reference nni provided :githublink:`pruner <nni/algorithms/compression/pytorch/pruning/one_shot_pruner.py>` implementations to implement your own pruner class.
----
-Customize a new quantization algorithm
--------------------------------------
-To write a new quantization algorithm, you can write a class that inherits ``nni.compression.pytorch.Quantizer``. Then, override the member functions with the logic of your algorithm. The member function to override is ``quantize_weight``. ``quantize_weight`` directly returns the quantized weights rather than mask, because for quantization the quantized weights cannot be obtained by applying mask.
-.. code-block:: python
-   from nni.compression.pytorch import Quantizer
-   class YourQuantizer(Quantizer):
-       def __init__(self, model, config_list):
-           """
-           Suggest you to use the NNI defined spec for config
-           """
-           super().__init__(model, config_list)
-       def quantize_weight(self, weight, config, **kwargs):
-           """
-           quantize should overload this method to quantize weight tensors.
-           This method is effectively hooked to :meth:`forward` of the model.
-           Parameters
-           ----------
-           weight : Tensor
-               weight that needs to be quantized
-           config : dict
-               the configuration for weight quantization
-           """
-           # Put your code to generate `new_weight` here
-           return new_weight
-       def quantize_output(self, output, config, **kwargs):
-           """
-           quantize should overload this method to quantize output.
-           This method is effectively hooked to `:meth:`forward` of the model.
-           Parameters
-           ----------
-           output : Tensor
-               output that needs to be quantized
-           config : dict
-               the configuration for output quantization
-           """
-           # Put your code to generate `new_output` here
-           return new_output
-       def quantize_input(self, *inputs, config, **kwargs):
-           """
-           quantize should overload this method to quantize input.
-           This method is effectively hooked to :meth:`forward` of the model.
-           Parameters
-           ----------
-           inputs : Tensor
-               inputs that needs to be quantized
-           config : dict
-               the configuration for inputs quantization
-           """
-           # Put your code to generate `new_input` here
-           return new_input
-       def update_epoch(self, epoch_num):
-           pass
-       def step(self):
-           """
-           Can do some processing based on the model or weights binded
-           in the func bind_model
-           """
-           pass
-Customize backward function
-^^^^^^^^^^^^^^^^^^^^^^^^^^^
-Sometimes it's necessary for a quantization operation to have a customized backward function, such as `Straight-Through Estimator <https://stackoverflow.com/questions/38361314/the-concept-of-straight-through-estimator-ste>`__ , user can customize a backward function as follow:
-.. code-block:: python
-   from nni.compression.pytorch.compressor import Quantizer, QuantGrad, QuantType
-   class ClipGrad(QuantGrad):
-       @staticmethod
-       def quant_backward(tensor, grad_output, quant_type):
-           """
-           This method should be overrided by subclass to provide customized backward function,
-           default implementation is Straight-Through Estimator
-           Parameters
-           ----------
-           tensor : Tensor
-               input of quantization operation
-           grad_output : Tensor
-               gradient of the output of quantization operation
-           quant_type : QuantType
-               the type of quantization, it can be `QuantType.INPUT`, `QuantType.WEIGHT`, `QuantType.OUTPUT`,
-               you can define different behavior for different types.
-           Returns
-           -------
-           tensor
-               gradient of the input of quantization operation
-           """
-           # for quant_output function, set grad to zero if the absolute value of tensor is larger than 1
-           if quant_type == QuantType.OUTPUT:
-               grad_output[torch.abs(tensor) > 1] = 0
-           return grad_output
-   class YourQuantizer(Quantizer):
-       def __init__(self, model, config_list):
-           super().__init__(model, config_list)
-           # set your customized backward function to overwrite default backward function
-           self.quant_grad = ClipGrad
-If you do not customize ``QuantGrad``\ , the default backward is Straight-Through Estimator. 
-*Coming Soon* ...
--- a/docs/source/compression/legacy_framework.rst
+++ b/docs/source/compression/legacy_framework.rst
-Framework overview of model compression
-=======================================
-.. contents::
-Below picture shows the components overview of model compression framework.
-.. image:: ../../img/compressor_framework.jpg
-   :target: ../../img/compressor_framework.jpg
-   :alt: 
-There are 3 major components/classes in NNI model compression framework: ``Compressor``\ , ``Pruner`` and ``Quantizer``. Let's look at them in detail one by one:
-Compressor
----------
-Compressor is the base class for pruner and quntizer, it provides a unified interface for pruner and quantizer for end users, so that pruner and quantizer can be used in the same way. For example, to use a pruner:
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import LevelPruner
-   # load a pretrained model or train a model before using a pruner
-   configure_list = [{
-       'sparsity': 0.7,
-       'op_types': ['Conv2d', 'Linear'],
-   }]
-   pruner = LevelPruner(model, configure_list)
-   model = pruner.compress()
-   # model is ready for pruning, now start finetune the model,
-   # the model will be pruned during training automatically
-To use a quantizer:
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import DoReFaQuantizer
-   configure_list = [{
-       'quant_types': ['weight'],
-       'quant_bits': {
-           'weight': 8,
-       },
-       'op_types':['Conv2d', 'Linear']
-   }]
-   optimizer = torch.optim.SGD(model.parameters(), lr=0.001, momentum=0.9, weight_decay=1e-4)
-   quantizer = DoReFaQuantizer(model, configure_list, optimizer)
-   quantizer.compress()
-View :githublink:`example code <examples/model_compress>` for more information.
-``Compressor`` class provides some utility methods for subclass and users:
-Set wrapper attribute
-^^^^^^^^^^^^^^^^^^^^^
-Sometimes ``calc_mask`` must save some state data, therefore users can use ``set_wrappers_attribute`` API to register attribute just like how buffers are registered in PyTorch modules. These buffers will be registered to ``module wrapper``. Users can access these buffers through ``module wrapper``.
-In above example, we use ``set_wrappers_attribute`` to set a buffer ``if_calculated`` which is used as flag indicating if the mask of a layer is already calculated.
-Collect data during forward
-^^^^^^^^^^^^^^^^^^^^^^^^^^^
-Sometimes users want to collect some data during the modules' forward method, for example, the mean value of the activation. This can be done by adding a customized collector to module.
-.. code-block:: python
-   class MyMasker(WeightMasker):
-       def __init__(self, model, pruner):
-           super().__init__(model, pruner)
-           # Set attribute `collected_activation` for all wrappers to store
-           # activations for each layer
-           self.pruner.set_wrappers_attribute("collected_activation", [])
-           self.activation = torch.nn.functional.relu
-           def collector(wrapper, input_, output):
-               # The collected activation can be accessed via each wrapper's collected_activation
-               # attribute
-               wrapper.collected_activation.append(self.activation(output.detach().cpu()))
-           self.pruner.hook_id = self.pruner.add_activation_collector(collector)
-The collector function will be called each time the forward method runs.
-Users can also remove this collector like this:
-.. code-block:: python
-   # Save the collector identifier
-   collector_id = self.pruner.add_activation_collector(collector)
-   # When the collector is not used any more, it can be remove using
-   # the saved collector identifier
-   self.pruner.remove_activation_collector(collector_id)
----
-Pruner
------
-A pruner receives ``model`` , ``config_list`` as arguments. 
-Some pruners like ``TaylorFOWeightFilter Pruner`` prune the model per the ``config_list`` during training loop by adding a hook on ``optimizer.step()``.
-Pruner class is a subclass of Compressor, so it contains everything in the Compressor class and some additional components only for pruning, it contains:
-Weight masker
-^^^^^^^^^^^^^
-A ``weight masker`` is the implementation of pruning algorithms, it can prune a specified layer wrapped by ``module wrapper`` with specified sparsity.
-Pruning module wrapper
-^^^^^^^^^^^^^^^^^^^^^^
-A ``pruning module wrapper`` is a module containing:
-#. the origin module
-#. some buffers used by ``calc_mask``
-#. a new forward method that applies masks before running the original forward method.
-the reasons to use ``module wrapper``\ :
-#. some buffers are needed by ``calc_mask`` to calculate masks and these buffers should be registered in ``module wrapper`` so that the original modules are not contaminated.
-#. a new ``forward`` method is needed to apply masks to weight before calling the real ``forward`` method.
-Pruning hook
-^^^^^^^^^^^^
-A pruning hook is installed on a pruner when the pruner is constructed, it is used to call pruner's calc_mask method at ``optimizer.step()`` is invoked.
----
-Quantizer
---------
-Quantizer class is also a subclass of ``Compressor``\ , it is used to compress models by reducing the number of bits required to represent weights or activations, which can reduce the computations and the inference time. It contains:
-Quantization module wrapper
-^^^^^^^^^^^^^^^^^^^^^^^^^^^
-Each module/layer of the model to be quantized is wrapped by a quantization module wrapper, it provides a new ``forward`` method to quantize the original module's weight, input and output.
-Quantization hook
-^^^^^^^^^^^^^^^^^
-A quantization hook is installed on a quntizer when it is constructed, it is call at ``optimizer.step()``.
-Quantization methods
-^^^^^^^^^^^^^^^^^^^^
-``Quantizer`` class provides following methods for subclass to implement quantization algorithms:
-.. code-block:: python
-   class Quantizer(Compressor):
-       """
-       Base quantizer for pytorch quantizer
-       """
-       def quantize_weight(self, weight, wrapper, **kwargs):
-           """
-           quantize should overload this method to quantize weight.
-           This method is effectively hooked to :meth:`forward` of the model.
-           Parameters
-           ----------
-           weight : Tensor
-               weight that needs to be quantized
-           wrapper : QuantizerModuleWrapper
-               the wrapper for origin module
-           """
-           raise NotImplementedError('Quantizer must overload quantize_weight()')
-       def quantize_output(self, output, wrapper, **kwargs):
-           """
-           quantize should overload this method to quantize output.
-           This method is effectively hooked to :meth:`forward` of the model.
-           Parameters
-           ----------
-           output : Tensor
-               output that needs to be quantized
-           wrapper : QuantizerModuleWrapper
-               the wrapper for origin module
-           """
-           raise NotImplementedError('Quantizer must overload quantize_output()')
-       def quantize_input(self, *inputs, wrapper, **kwargs):
-           """
-           quantize should overload this method to quantize input.
-           This method is effectively hooked to :meth:`forward` of the model.
-           Parameters
-           ----------
-           inputs : Tensor
-               inputs that needs to be quantized
-           wrapper : QuantizerModuleWrapper
-               the wrapper for origin module
-           """
-           raise NotImplementedError('Quantizer must overload quantize_input()')
----
-Multi-GPU support
-----------------
-On multi-GPU training, buffers and parameters are copied to multiple GPU every time the ``forward`` method runs on multiple GPU. If buffers and parameters are updated in the ``forward`` method, an ``in-place`` update is needed to ensure the update is effective.
-Since ``calc_mask`` is called in the ``optimizer.step`` method, which happens after the ``forward`` method and happens only on one GPU, it supports multi-GPU naturally.
--- a/docs/source/reference/pruner.rst
+++ b/docs/source/reference/pruner.rst
--- a/docs/source/compression/pruning.rst
+++ b/docs/source/compression/pruning.rst
 Model Pruning with NNI
 ======================
-Pruning V2 is a refactoring of the old version and provides more powerful functions.
+Pruning is a common technique to compress neural network models.
-Compared with the old version, the iterative pruning process is detached from the pruner and the pruner is only responsible for pruning and generating the masks once.
+The pruning methods explore the redundancy in the model weights(parameters) and try to remove/prune the redundant and uncritical weights.
-What's more, pruning V2 unifies the pruning process and provides a more free combination of pruning components.
+The redundant elements are pruned from the model, their values are zeroed and we make sure they don't take part in the back-propagation process.
-Task generator only cares about the pruning effect that should be achieved in each round, and uses a config list to express how to pruning in the next step.
-Pruner will reset with the model and config list given by task generator then generate the masks in current step.
+The following concepts can help you understand pruning in NNI.
+.. Using rubric to prevent the section heading to be include into toc
+.. rubric:: Pruning Target
+Pruning target means where we apply the sparsity.
+Most pruning methods prune the weights to reduce the model size and accelerate the inference latency.
+Other pruning methods also apply sparsity on the inputs, outputs or intermediate states to accelerate the inference latency.
+NNI support pruning module weights right now, and will support other pruning targets in the future.
+.. rubric:: Basic Pruner
+Basic pruner generates the masks for each pruning targets (weights) for a determined sparsity ratio.
+It usually takes model and config as input arguments, then generate a mask for the model.
+.. rubric:: Scheduled Pruner
+Scheduled pruner decides how to allocate sparsity ratio to each pruning targets, it also handles the pruning speed up and finetuning logic.
+From the implementation logic, the scheduled pruner is a combination of pruning scheduler, basic pruner and task generator.
+Task generator only cares about the pruning effect that should be achieved in each round, and uses a config list to express how to pruning.
+Basic pruner will reset with the model and config list given by task generator then generate the masks.
 For a clearer structure vision, please refer to the figure below.
 .. image:: ../../img/pruning_process.png
   :target: ../../img/pruning_process.png
+   :scale: 80%
+   :align: center
   :alt:
-A pruning process is usually driven by a pruning scheduler, it contains a specific pruner and a task generator.
+More information about scheduled pruning process please refer to :doc:`Pruning Scheduler <pruning_scheduler>`.
+.. rubric:: Granularity
+Fine-grained pruning or unstructured pruning refers to pruning each individual weights separately.
+Coarse-grained pruning or structured pruning is pruning entire group of weights, such as a convolutional filter.
+:ref:`level-pruner` is the only fine-grained pruner in NNI, all other pruners pruning the output channels on weights.
+.. _dependency-awareode-for-output-channel-pruning:
+.. rubric:: Dependency-aware Mode for Output Channel Pruning
+Currently, we support ``dependency aware`` mode in several ``pruner``: :ref:`l1-norm-pruner`, :ref:`l2-norm-pruner`, :ref:`fpgm-pruner`,
+:ref:`activation-apoz-rank-pruner`, :ref:`activation-mean-rank-pruner`, :ref:`taylor-fo-weight-pruner`.
+In these pruning algorithms, the pruner will prune each layer separately. While pruning a layer,
+the algorithm will quantify the importance of each filter based on some specific rules(such as l1 norm), and prune the less important output channels.
+We use pruning convolutional layers as an example to explain ``dependency aware`` mode.
+As :doc:`dependency analysis utils <./compression_utils>` shows, if the output channels of two convolutional layers(conv1, conv2) are added together,
+then these two convolutional layers have channel dependency with each other(more details please see :doc:`Compression Utils <./compression_utils>` ).
+Take the following figure as an example.
+.. image:: ../../img/mask_conflict.jpg
+   :target: ../../img/mask_conflict.jpg
+   :scale: 80%
+   :align: center
+   :alt: 
+If we prune the first 50% of output channels (filters) for conv1, and prune the last 50% of output channels for conv2.
+Although both layers have pruned 50% of the filters, the speedup module still needs to add zeros to align the output channels.
+In this case, we cannot harvest the speed benefit from the model pruning.
+To better gain the speed benefit of the model pruning, we add a dependency-aware mode for the ``Pruner`` that can prune the output channels.
+In the dependency-aware mode, the pruner prunes the model not only based on the metric of each output channels, but also the topology of the whole network architecture.
+In the dependency-aware mode (``dependency_aware`` is set ``True``), the pruner will try to prune the same output channels for the layers that have the channel dependencies with each other, as shown in the following figure.
+.. image:: ../../img/dependency-aware.jpg
+   :target: ../../img/dependency-aware.jpg
+   :scale: 80%
+   :align: center
+   :alt: 
+Take the dependency-aware mode of :ref:`l1-norm-pruner` as an example.
+Specifically, the pruner will calculate the L1 norm (for example) sum of all the layers in the dependency set for each channel.
+Obviously, the number of channels that can actually be pruned of this dependency set in the end is determined by the minimum sparsity of layers in this dependency set (denoted by ``min_sparsity``).
+According to the L1 norm sum of each channel, the pruner will prune the same ``min_sparsity`` channels for all the layers.
+Next, the pruner will additionally prune ``sparsity`` - ``min_sparsity`` channels for each convolutional layer based on its own L1 norm of each channel.
+For example, suppose the output channels of ``conv1``, ``conv2`` are added together and the configured sparsities of ``conv1`` and ``conv2`` are 0.3, 0.2 respectively.
+In this case, the ``dependency-aware pruner`` will 
+* First, prune the same 20% of channels for `conv1` and `conv2` according to L1 norm sum of `conv1` and `conv2`.
+* Second, the pruner will additionally prune 10% channels for `conv1` according to the L1 norm of each channel of `conv1`.
-.. Note::
+In addition, for the convolutional layers that have more than one filter group,
+``dependency-aware pruner`` will also try to prune the same number of the channels for each filter group.
+Overall, this pruner will prune the model according to the L1 norm of each filter and try to meet the topological constrains (channel dependency, etc) to improve the final speed gain after the speedup process. 
-    But users can also use pruner directly like in the pruning V1.
+In the dependency-aware mode, the pruner will provide a better speed gain from the model pruning.
-..  toctree::
+.. toctree::
+    :hidden:
    :maxdepth: 2
    Quickstart <../tutorials/pruning_quick_start_mnist>
-    Concepts <pruning_concepts>
+    Pruner Reference <pruner>
    Speed Up <../tutorials/pruning_speed_up>
-    Pruner V2 Reference <../reference/pruner>
-    Pruner Reference (legacy) <../reference/legacy_pruner>
--- a/docs/source/compression/pruning_concepts.rst
+++ b/docs/source/compression/pruning_concepts.rst
-Pruning Concepts
-================
-Pruning is a common technique to compress neural network models.
-The pruning methods explore the redundancy in the model weights(parameters) and try to remove/prune the redundant and uncritical weights.
-The redundant elements are pruned from the model, their values are zeroed and we make sure they don't take part in the back-propagation process.
-In NNI, a pruning method is divided into multiple dimensions.
-Pruning Target
--------------
-Pruning target means where we apply the sparsity.
-Most pruning methods prune the weight to reduce the model size and accelerate the inference latency.
-Other pruning methods also apply sparsity on the input and output to accelerate the inference latency.
-NNI support pruning module weight right now, and will support pruning input & output in the future.
-Basic Pruners & Scheduled Pruners
---------------------------------
-Basic pruners generate the masks for each pruning targets (weights) for a determined sparsity ratio.
-Scheduled pruners decide how to allocate sparsity ratio to each pruning targets, they always work with basic pruner to generate masks.
-Granularity
-----------
-Fine-grained pruning or unstructured pruning refers to pruning each individual weights separately.
-Coarse-grained pruning or structured pruning is pruning entire group of weights, such as a convolutional filter.
-:ref:`level-pruner` is the only fine-grained pruner in NNI, all other pruners pruning the output channels on weights.
-.. _dependency-awareode-for-output-channel-pruning:
-Dependency-aware Mode for Output Channel Pruning
------------------------------------------------
-Currently, we support ``dependency aware`` mode in several ``pruner``: :ref:`l1-norm-pruner`, :ref:`l2-norm-pruner`, :ref:`fpgm-pruner`,
-:ref:`activation-apoz-rank-pruner`, :ref:`activation-mean-rank-pruner`, :ref:`taylor-fo-weight-pruner`.
-In these pruning algorithms, the pruner will prune each layer separately. While pruning a layer,
-the algorithm will quantify the importance of each filter based on some specific rules(such as l1 norm), and prune the less important output channels.
-We use pruning convolutional layers as an example to explain ``dependency aware`` mode.
-As :doc:`dependency analysis utils <./compression_utils>` shows, if the output channels of two convolutional layers(conv1, conv2) are added together,
-then these two convolutional layers have channel dependency with each other(more details please see :doc:`Compression Utils <./compression_utils>` ).
-Take the following figure as an example.
-.. image:: ../../img/mask_conflict.jpg
-   :target: ../../img/mask_conflict.jpg
-   :alt: 
-If we prune the first 50% of output channels (filters) for conv1, and prune the last 50% of output channels for conv2.
-Although both layers have pruned 50% of the filters, the speedup module still needs to add zeros to align the output channels.
-In this case, we cannot harvest the speed benefit from the model pruning.
-To better gain the speed benefit of the model pruning, we add a dependency-aware mode for the ``Pruner`` that can prune the output channels.
-In the dependency-aware mode, the pruner prunes the model not only based on the metric of each output channels, but also the topology of the whole network architecture.
-In the dependency-aware mode (``dependency_aware`` is set ``True``), the pruner will try to prune the same output channels for the layers that have the channel dependencies with each other, as shown in the following figure.
-.. image:: ../../img/dependency-aware.jpg
-   :target: ../../img/dependency-aware.jpg
-   :alt: 
-Take the dependency-aware mode of :ref:`l1-norm-pruner` as an example.
-Specifically, the pruner will calculate the L1 norm (for example) sum of all the layers in the dependency set for each channel.
-Obviously, the number of channels that can actually be pruned of this dependency set in the end is determined by the minimum sparsity of layers in this dependency set (denoted by ``min_sparsity``).
-According to the L1 norm sum of each channel, the pruner will prune the same ``min_sparsity`` channels for all the layers.
-Next, the pruner will additionally prune ``sparsity`` - ``min_sparsity`` channels for each convolutional layer based on its own L1 norm of each channel.
-For example, suppose the output channels of ``conv1``, ``conv2`` are added together and the configured sparsities of ``conv1`` and ``conv2`` are 0.3, 0.2 respectively.
-In this case, the ``dependency-aware pruner`` will 
-* First, prune the same 20% of channels for `conv1` and `conv2` according to L1 norm sum of `conv1` and `conv2`.
-* Second, the pruner will additionally prune 10% channels for `conv1` according to the L1 norm of each channel of `conv1`.
-In addition, for the convolutional layers that have more than one filter group,
-``dependency-aware pruner`` will also try to prune the same number of the channels for each filter group.
-Overall, this pruner will prune the model according to the L1 norm of each filter and try to meet the topological constrains (channel dependency, etc) to improve the final speed gain after the speedup process. 
-In the dependency-aware mode, the pruner will provide a better speed gain from the model pruning.
--- a/docs/source/compression/quantization.rst
+++ b/docs/source/compression/quantization.rst
 Model Quantization with NNI
 ===========================
-..  toctree::
+Quantization refers to compressing models by reducing the number of bits required to represent weights or activations,
+which can reduce the computations and the inference time. In the context of deep neural networks, the major numerical
+format for model weights is 32-bit float, or FP32. Many research works have demonstrated that weights and activations
+can be represented using 8-bit integers without significant loss in accuracy. Even lower bit-widths, such as 4/2/1 bits,
+is an active field of research.
+A quantizer is a quantization algorithm implementation in NNI, NNI provides multiple quantizers as below. You can also
+create your own quantizer using NNI model compression interface.
+.. toctree::
+    :hidden:
    :maxdepth: 2
    Quickstart <../tutorials/quantization_quick_start_mnist>
+    Quantizer Reference <quantizer>
    Speed Up <../tutorials/quantization_speed_up>
-    Quantizer Reference <../reference/quantizer>
--- a/docs/source/reference/quantizer.rst
+++ b/docs/source/reference/quantizer.rst
--- a/docs/source/compression/reference.rst
+++ b/docs/source/compression/reference.rst
-Reference
-=========
-..  toctree::
-    :maxdepth: 2
-    Pruner V2 Reference <../reference/pruner>
-    Pruner Reference (legacy) <../reference/legacy_pruner>
-    Quantizer Reference <../reference/quantizer>
-    Compression Config Specification <../reference/compression_config_list>
--- a/docs/source/index.rst
+++ b/docs/source/index.rst
@@ -23,7 +23,7 @@ Neural Network Intelligence
    Overview
    Auto (Hyper-parameter) Tuning <hyperparameter_tune>
    Neural Architecture Search <nas/index>
-    Model Compression <compression>
+    Model Compression <compression/index>
    Feature Engineering <feature_engineering>
    Experiment <experiment/overview>

--- a/docs/source/index_zh.rst
+++ b/docs/source/index_zh.rst
-.. cbe5c6f0f6b5a054dc36d05b49d1986e
+.. 84633d9c4ebf3421e7618c56117045c2
 ###########################
 Neural Network Intelligence
@@ -16,7 +16,7 @@ Neural Network Intelligence
    教程<tutorials>
    自动（超参数）调优 <hyperparameter_tune>
    神经网络架构搜索<nas/index>
-    模型压缩<compression>
+    模型压缩<compression/index>
    特征工程<feature_engineering>
    NNI实验 <experiment/overview>
    参考<reference>

--- a/docs/source/reference.rst
+++ b/docs/source/reference.rst
@@ -17,4 +17,3 @@ References
    Supported Framework Library <SupportedFramework_Library>
    Launch from Python <Tutorial/HowToLaunchFromPython>
    Tensorboard <Tutorial/Tensorboard>
-    Compression <compression/reference>
--- a/docs/source/reference/legacy_pruner.rst
+++ b/docs/source/reference/legacy_pruner.rst
-Supported Pruning Algorithms on NNI
-===================================
-We provide several pruning algorithms that support fine-grained weight pruning and structural filter pruning. **Fine-grained Pruning** generally results in  unstructured models, which need specialized hardware or software to speed up the sparse network. **Filter Pruning** achieves acceleration by removing the entire filter. Some pruning algorithms use one-shot method that prune weights at once based on an importance metric (It is necessary to finetune the model to compensate for the loss of accuracy). Other pruning algorithms **iteratively** prune weights during optimization, which control the pruning schedule, including some automatic pruning algorithms.
-**One-shot Pruning**
-* `Level Pruner <#level-pruner>`__ ((fine-grained pruning))
-* `Slim Pruner <#slim-pruner>`__
-* `FPGM Pruner <#fpgm-pruner>`__
-* `L1Filter Pruner <#l1filter-pruner>`__
-* `L2Filter Pruner <#l2filter-pruner>`__
-* `Activation APoZ Rank Filter Pruner <#activationAPoZRankFilter-pruner>`__
-* `Activation Mean Rank Filter Pruner <#activationmeanrankfilter-pruner>`__
-* `Taylor FO On Weight Pruner <#taylorfoweightfilter-pruner>`__
-**Iteratively Pruning**
-* `AGP Pruner <#agp-pruner>`__
-* `NetAdapt Pruner <#netadapt-pruner>`__
-* `SimulatedAnnealing Pruner <#simulatedannealing-pruner>`__
-* `AutoCompress Pruner <#autocompress-pruner>`__
-* `AMC Pruner <#amc-pruner>`__
-* `Sensitivity Pruner <#sensitivity-pruner>`__
-* `ADMM Pruner <#admm-pruner>`__
-**Others**
-* `Lottery Ticket Hypothesis <#lottery-ticket-hypothesis>`__
-* `Transformer Head Pruner <#transformer-head-pruner>`__
-Level Pruner
------------
-This is one basic one-shot pruner: you can set a target sparsity level (expressed as a fraction, 0.6 means we will prune 60% of the weight parameters). 
-We first sort the weights in the specified layer by their absolute values. And then mask to zero the smallest magnitude weights until the desired sparsity level is reached.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import LevelPruner
-   config_list = [{ 'sparsity': 0.8, 'op_types': ['default'] }]
-   pruner = LevelPruner(model, config_list)
-   pruner.compress()
-User configuration for Level Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.LevelPruner
-**TensorFlow**
-..  autoclass:: nni.algorithms.compression.tensorflow.pruning.LevelPruner
-Slim Pruner
-----------
-This is an one-shot pruner, which adds sparsity regularization on the scaling factors of batch normalization (BN) layers during training to identify unimportant channels. The channels with small scaling factor values will be pruned. For more details, please refer to `'Learning Efficient Convolutional Networks through Network Slimming' <https://arxiv.org/pdf/1708.06519.pdf>`__\.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import SlimPruner
-   config_list = [{ 'sparsity': 0.8, 'op_types': ['BatchNorm2d'] }]
-   pruner = SlimPruner(model, config_list, optimizer, trainer, criterion)
-   pruner.compress()
-User configuration for Slim Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.SlimPruner
-Reproduced Experiment
-^^^^^^^^^^^^^^^^^^^^^
-We implemented one of the experiments in `Learning Efficient Convolutional Networks through Network Slimming <https://arxiv.org/pdf/1708.06519.pdf>`__\ , we pruned ``70%`` channels in the **VGGNet** for CIFAR-10 in the paper, in which ``88.5%`` parameters are pruned. Our experiments results are as follows:
-.. list-table::
-   :header-rows: 1
-   :widths: auto
-   * - Model
-     - Error(paper/ours)
-     - Parameters
-     - Pruned
-   * - VGGNet
-     - 6.34/6.69
-     - 20.04M
-     - 
-   * - Pruned-VGGNet
-     - 6.20/6.34
-     - 2.03M
-     - 88.5%
-The experiments code can be found at :githublink:`examples/model_compress/pruning/basic_pruners_torch.py <examples/model_compress/pruning/basic_pruners_torch.py>`
-.. code-block:: python
-   python basic_pruners_torch.py --pruner slim --model vgg19 --sparsity 0.7 --speed-up
----
-FPGM Pruner
-----------
-This is an one-shot pruner, which prunes filters with the smallest geometric median. FPGM chooses the filters with the most replaceable contribution.
-For more details, please refer to `Filter Pruning via Geometric Median for Deep Convolutional Neural Networks Acceleration <https://arxiv.org/pdf/1811.00250.pdf>`__.
-We also provide a dependency-aware mode for this pruner to get better speedup from the pruning. Please reference :ref:`dependency-awareode-for-output-channel-pruning` for more details.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import FPGMPruner
-   config_list = [{
-       'sparsity': 0.5,
-       'op_types': ['Conv2d']
-   }]
-   pruner = FPGMPruner(model, config_list)
-   pruner.compress()
-User configuration for FPGM Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.FPGMPruner
-L1Filter Pruner
---------------
-This is an one-shot pruner, which prunes the filters in the **convolution layers**.
-..
-   The procedure of pruning m filters from the ith convolutional layer is as follows:
-   #. For each filter :math:`F_{i,j}`, calculate the sum of its absolute kernel weights :math:`s_j=\sum_{l=1}^{n_i}\sum|K_l|`.
-   #. Sort the filters by :math:`s_j`.
-   #. Prune :math:`m` filters with the smallest sum values and their corresponding feature maps. The
-      kernels in the next convolutional layer corresponding to the pruned feature maps are also removed.
-   #. A new kernel matrix is created for both the :math:`i`-th and :math:`i+1`-th layers, and the remaining kernel
-      weights are copied to the new model.
-For more details, please refer to `PRUNING FILTERS FOR EFFICIENT CONVNETS <https://arxiv.org/abs/1608.08710>`__\.
-In addition, we also provide a dependency-aware mode for the L1FilterPruner. For more details about the dependency-aware mode, please reference :ref:`dependency-awareode-for-output-channel-pruning`.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import L1FilterPruner
-   config_list = [{ 'sparsity': 0.8, 'op_types': ['Conv2d'] }]
-   pruner = L1FilterPruner(model, config_list)
-   pruner.compress()
-User configuration for L1Filter Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.L1FilterPruner
-Reproduced Experiment
-^^^^^^^^^^^^^^^^^^^^^
-We implemented one of the experiments in `PRUNING FILTERS FOR EFFICIENT CONVNETS <https://arxiv.org/abs/1608.08710>`__ with **L1FilterPruner**\ , we pruned **VGG-16** for CIFAR-10 to **VGG-16-pruned-A** in the paper, in which ``64%`` parameters are pruned. Our experiments results are as follows:
-.. list-table::
-   :header-rows: 1
-   :widths: auto
-   * - Model
-     - Error(paper/ours)
-     - Parameters
-     - Pruned
-   * - VGG-16
-     - 6.75/6.49
-     - 1.5x10^7
-     - 
-   * - VGG-16-pruned-A
-     - 6.60/6.47
-     - 5.4x10^6
-     - 64.0%
-The experiments code can be found at :githublink:`examples/model_compress/pruning/basic_pruners_torch.py <examples/model_compress/pruning/basic_pruners_torch.py>`
-.. code-block:: python
-   python basic_pruners_torch.py --pruner l1filter --model vgg16 --speed-up
----
-L2Filter Pruner
---------------
-This is a structured pruning algorithm that prunes the filters with the smallest L2 norm of the weights. It is implemented as a one-shot pruner.
-We also provide a dependency-aware mode for this pruner to get better speedup from the pruning. Please reference :ref:`dependency-awareode-for-output-channel-pruning` for more details.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import L2FilterPruner
-   config_list = [{ 'sparsity': 0.8, 'op_types': ['Conv2d'] }]
-   pruner = L2FilterPruner(model, config_list)
-   pruner.compress()
-User configuration for L2Filter Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.L2FilterPruner
----
-ActivationAPoZRankFilter Pruner
-------------------------------
-ActivationAPoZRankFilter Pruner is a pruner which prunes the filters with the smallest importance criterion ``APoZ`` calculated from the output activations of convolution layers to achieve a preset level of network sparsity. The pruning criterion ``APoZ`` is explained in the paper `Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures <https://arxiv.org/abs/1607.03250>`__.
-The APoZ is defined as:
-:math:`APoZ_{c}^{(i)} = APoZ\left(O_{c}^{(i)}\right)=\frac{\sum_{k}^{N} \sum_{j}^{M} f\left(O_{c, j}^{(i)}(k)=0\right)}{N \times M}`
-We also provide a dependency-aware mode for this pruner to get better speedup from the pruning. Please reference :ref:`dependency-awareode-for-output-channel-pruning` for more details.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import ActivationAPoZRankFilterPruner
-   config_list = [{
-       'sparsity': 0.5,
-       'op_types': ['Conv2d']
-   }]
-   pruner = ActivationAPoZRankFilterPruner(model, config_list, optimizer, trainer, criterion, sparsifying_training_batches=1)
-   pruner.compress()
-Note: ActivationAPoZRankFilterPruner is used to prune convolutional layers within deep neural networks, therefore the ``op_types`` field supports only convolutional layers.
-You can view :githublink:`example <examples/model_compress/pruning/basic_pruners_torch.py>` for more information.
-User configuration for ActivationAPoZRankFilter Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.ActivationAPoZRankFilterPruner
----
-ActivationMeanRankFilter Pruner
-------------------------------
-ActivationMeanRankFilterPruner is a pruner which prunes the filters with the smallest importance criterion ``mean activation`` calculated from the output activations of convolution layers to achieve a preset level of network sparsity. The pruning criterion ``mean activation`` is explained in section 2.2 of the paper `Pruning Convolutional Neural Networks for Resource Efficient Inference <https://arxiv.org/abs/1611.06440>`__. Other pruning criteria mentioned in this paper will be supported in future release.
-We also provide a dependency-aware mode for this pruner to get better speedup from the pruning. Please reference :ref:`dependency-awareode-for-output-channel-pruning` for more details.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import ActivationMeanRankFilterPruner
-   config_list = [{
-       'sparsity': 0.5,
-       'op_types': ['Conv2d']
-   }]
-   pruner = ActivationMeanRankFilterPruner(model, config_list, optimizer, trainer, criterion, sparsifying_training_batches=1)
-   pruner.compress()
-Note: ActivationMeanRankFilterPruner is used to prune convolutional layers within deep neural networks, therefore the ``op_types`` field supports only convolutional layers.
-You can view :githublink:`example <examples/model_compress/pruning/basic_pruners_torch.py>` for more information.
-User configuration for ActivationMeanRankFilterPruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.ActivationMeanRankFilterPruner
----
-TaylorFOWeightFilter Pruner
---------------------------
-TaylorFOWeightFilter Pruner is a pruner which prunes convolutional layers based on estimated importance calculated from the first order taylor expansion on weights to achieve a preset level of network sparsity. The estimated importance of filters is defined as the paper `Importance Estimation for Neural Network Pruning <http://jankautz.com/publications/Importance4NNPruning_CVPR19.pdf>`__. Other pruning criteria mentioned in this paper will be supported in future release.
-..
-:math:`\widehat{\mathcal{I}}_{\mathcal{S}}^{(1)}(\mathbf{W}) \triangleq \sum_{s \in \mathcal{S}} \mathcal{I}_{s}^{(1)}(\mathbf{W})=\sum_{s \in \mathcal{S}}\left(g_{s} w_{s}\right)^{2}`
-We also provide a dependency-aware mode for this pruner to get better speedup from the pruning. Please reference :ref:`dependency-awareode-for-output-channel-pruning` for more details.
-What's more, we provide a global-sort mode for this pruner which is aligned with paper implementation. Please set parameter 'global_sort' to True when instantiate TaylorFOWeightFilterPruner.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import TaylorFOWeightFilterPruner
-   config_list = [{
-       'sparsity': 0.5,
-       'op_types': ['Conv2d']
-   }]
-   pruner = TaylorFOWeightFilterPruner(model, config_list, optimizer, trainer, criterion, sparsifying_training_batches=1)
-   pruner.compress()
-User configuration for TaylorFOWeightFilter Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.TaylorFOWeightFilterPruner
----
-AGP Pruner
----------
-This is an iterative pruner, which the sparsity is increased from an initial sparsity value si (usually 0) to a final sparsity value sf over a span of n pruning steps, starting at training step :math:`t_{0}` and with pruning frequency :math:`\Delta t`:
-:math:`s_{t}=s_{f}+\left(s_{i}-s_{f}\right)\left(1-\frac{t-t_{0}}{n \Delta t}\right)^{3} \text { for } t \in\left\{t_{0}, t_{0}+\Delta t, \ldots, t_{0} + n \Delta t\right\}`
-For more details please refer to `To prune, or not to prune: exploring the efficacy of pruning for model compression <https://arxiv.org/abs/1710.01878>`__\.
-Usage
-^^^^^
-You can prune all weights from 0% to 80% sparsity in 10 epoch with the code below.
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import AGPPruner
-   config_list = [{
-       'sparsity': 0.8,
-       'op_types': ['default']
-   }]
-   # load a pretrained model or train a model before using a pruner
-   # model = MyModel()
-   # model.load_state_dict(torch.load('mycheckpoint.pth'))
-   # AGP pruner prunes model while fine tuning the model by adding a hook on
-   # optimizer.step(), so an optimizer is required to prune the model.
-   optimizer = torch.optim.SGD(model.parameters(), lr=0.001, momentum=0.9, weight_decay=1e-4)
-   pruner = AGPPruner(model, config_list, optimizer, trainer, criterion, pruning_algorithm='level')
-   pruner.compress()
-AGP pruner uses ``LevelPruner`` algorithms to prune the weight by default, however you can set ``pruning_algorithm`` parameter to other values to use other pruning algorithms:
-* ``level``\ : LevelPruner
-* ``slim``\ : SlimPruner
-* ``l1``\ : L1FilterPruner
-* ``l2``\ : L2FilterPruner
-* ``fpgm``\ : FPGMPruner
-* ``taylorfo``\ : TaylorFOWeightFilterPruner
-* ``apoz``\ : ActivationAPoZRankFilterPruner
-* ``mean_activation``\ : ActivationMeanRankFilterPruner
-User configuration for AGP Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.AGPPruner
----
-NetAdapt Pruner
---------------
-NetAdapt allows a user to automatically simplify a pretrained network to meet the resource budget. 
-Given the overall sparsity, NetAdapt will automatically generate the sparsities distribution among different layers by iterative pruning.
-For more details, please refer to `NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications <https://arxiv.org/abs/1804.03230>`__.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import NetAdaptPruner
-   config_list = [{
-       'sparsity': 0.5,
-       'op_types': ['Conv2d']
-   }]
-   pruner = NetAdaptPruner(model, config_list, short_term_fine_tuner=short_term_fine_tuner, evaluator=evaluator,base_algo='l1', experiment_data_dir='./')
-   pruner.compress()
-You can view :githublink:`example <examples/model_compress/pruning/auto_pruners_torch.py>` for more information.
-User configuration for NetAdapt Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.NetAdaptPruner
-SimulatedAnnealing Pruner
-------------------------
-We implement a guided heuristic search method, Simulated Annealing (SA) algorithm, with enhancement on guided search based on prior experience. 
-The enhanced SA technique is based on the observation that a DNN layer with more number of weights often has a higher degree of model compression with less impact on overall accuracy.
-* Randomly initialize a pruning rate distribution (sparsities).
-* While current_temperature < stop_temperature:
-  #. generate a perturbation to current distribution
-  #. Perform fast evaluation on the perturbated distribution
-  #. accept the perturbation according to the performance and probability, if not accepted, return to step 1
-  #. cool down, current_temperature <- current_temperature * cool_down_rate
-For more details, please refer to `AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates <https://arxiv.org/abs/1907.03141>`__.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import SimulatedAnnealingPruner
-   config_list = [{
-       'sparsity': 0.5,
-       'op_types': ['Conv2d']
-   }]
-   pruner = SimulatedAnnealingPruner(model, config_list, evaluator=evaluator, base_algo='l1', cool_down_rate=0.9, experiment_data_dir='./')
-   pruner.compress()
-You can view :githublink:`example <examples/model_compress/pruning/auto_pruners_torch.py>` for more information.
-User configuration for SimulatedAnnealing Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.SimulatedAnnealingPruner
-AutoCompress Pruner
-------------------
-For each round, AutoCompressPruner prune the model for the same sparsity to achive the overall sparsity:
-.. code-block:: bash
-       1. Generate sparsities distribution using SimulatedAnnealingPruner
-       2. Perform ADMM-based structured pruning to generate pruning result for the next round.
-          Here we use `speedup` to perform real pruning.
-For more details, please refer to `AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates <https://arxiv.org/abs/1907.03141>`__.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import AutoCompressPruner
-   config_list = [{
-           'sparsity': 0.5,
-           'op_types': ['Conv2d']
-       }]
-   pruner = AutoCompressPruner(
-               model, config_list, trainer=trainer, evaluator=evaluator,
-               dummy_input=dummy_input, num_iterations=3, optimize_mode='maximize', base_algo='l1',
-               cool_down_rate=0.9, admm_num_iterations=30, admm_training_epochs=5, experiment_data_dir='./')
-   pruner.compress()
-You can view :githublink:`example <examples/model_compress/pruning/auto_pruners_torch.py>` for more information.
-User configuration for AutoCompress Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.AutoCompressPruner
-AMC Pruner
----------
-AMC pruner leverages reinforcement learning to provide the model compression policy.
-This learning-based compression policy outperforms conventional rule-based compression policy by having higher compression ratio,
-better preserving the accuracy and freeing human labor.
-For more details, please refer to `AMC: AutoML for Model Compression and Acceleration on Mobile Devices <https://arxiv.org/pdf/1802.03494.pdf>`__.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import AMCPruner
-   config_list = [{
-           'op_types': ['Conv2d', 'Linear']
-       }]
-   pruner = AMCPruner(model, config_list, evaluator, val_loader, flops_ratio=0.5)
-   pruner.compress()
-You can view :githublink:`example <examples/model_compress/pruning/amc/>` for more information.
-User configuration for AMC Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.AMCPruner
-Reproduced Experiment
-^^^^^^^^^^^^^^^^^^^^^
-We implemented one of the experiments in `AMC: AutoML for Model Compression and Acceleration on Mobile Devices <https://arxiv.org/pdf/1802.03494.pdf>`__\ , we pruned **MobileNet** to 50% FLOPS for ImageNet in the paper. Our experiments results are as follows:
-.. list-table::
-   :header-rows: 1
-   :widths: auto
-   * - Model
-     - Top 1 acc.(paper/ours)
-     - Top 5 acc. (paper/ours)
-     - FLOPS
-   * - MobileNet
-     - 70.5% / 69.9%
-     - 89.3% / 89.1%
-     - 50%
-The experiments code can be found at :githublink:`examples/model_compress/pruning/ <examples/model_compress/pruning/amc/>`
-ADMM Pruner
-----------
-Alternating Direction Method of Multipliers (ADMM) is a mathematical optimization technique,
-by decomposing the original nonconvex problem into two subproblems that can be solved iteratively. In weight pruning problem, these two subproblems are solved via 1) gradient descent algorithm and 2) Euclidean projection respectively. 
-During the process of solving these two subproblems, the weights of the original model will be changed. An one-shot pruner will then be applied to prune the model according to the config list given.
-This solution framework applies both to non-structured and different variations of structured pruning schemes.
-For more details, please refer to `A Systematic DNN Weight Pruning Framework using Alternating Direction Method of Multipliers <https://arxiv.org/abs/1804.03294>`__.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import ADMMPruner
-   config_list = [{
-               'sparsity': 0.8,
-               'op_types': ['Conv2d'],
-               'op_names': ['conv1']
-           }, {
-               'sparsity': 0.92,
-               'op_types': ['Conv2d'],
-               'op_names': ['conv2']
-           }]
-   pruner = ADMMPruner(model, config_list, trainer, num_iterations=30, epochs_per_iteration=5)
-   pruner.compress()
-You can view :githublink:`example <examples/model_compress/pruning/auto_pruners_torch.py>` for more information.
-User configuration for ADMM Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.ADMMPruner
-Lottery Ticket Hypothesis
-------------------------
-`The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks <https://arxiv.org/abs/1803.03635>`__\ , authors Jonathan Frankle and Michael Carbin,provides comprehensive measurement and analysis, and articulate the *lottery ticket hypothesis*\ : dense, randomly-initialized, feed-forward networks contain subnetworks (*winning tickets*\ ) that -- when trained in isolation -- reach test accuracy comparable to the original network in a similar number of iterations.
-In this paper, the authors use the following process to prune a model, called *iterative prunning*\ :
-..
-   #. Randomly initialize a neural network f(x;theta_0) (where theta\ *0 follows D*\ {theta}).
-   #. Train the network for j iterations, arriving at parameters theta_j.
-   #. Prune p% of the parameters in theta_j, creating a mask m.
-   #. Reset the remaining parameters to their values in theta_0, creating the winning ticket f(x;m*theta_0).
-   #. Repeat step 2, 3, and 4.
-If the configured final sparsity is P (e.g., 0.8) and there are n times iterative pruning, each iterative pruning prunes 1-(1-P)^(1/n) of the weights that survive the previous round.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import LotteryTicketPruner
-   config_list = [{
-       'prune_iterations': 5,
-       'sparsity': 0.8,
-       'op_types': ['default']
-   }]
-   pruner = LotteryTicketPruner(model, config_list, optimizer)
-   pruner.compress()
-   for _ in pruner.get_prune_iterations():
-       pruner.prune_iteration_start()
-       for epoch in range(epoch_num):
-           ...
-The above configuration means that there are 5 times of iterative pruning. As the 5 times iterative pruning are executed in the same run, LotteryTicketPruner needs ``model`` and ``optimizer`` (\ **Note that should add ``lr_scheduler`` if used**\ ) to reset their states every time a new prune iteration starts. Please use ``get_prune_iterations`` to get the pruning iterations, and invoke ``prune_iteration_start`` at the beginning of each iteration. ``epoch_num`` is better to be large enough for model convergence, because the hypothesis is that the performance (accuracy) got in latter rounds with high sparsity could be comparable with that got in the first round.
-User configuration for LotteryTicket Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.LotteryTicketPruner
-Reproduced Experiment
-^^^^^^^^^^^^^^^^^^^^^
-We try to reproduce the experiment result of the fully connected network on MNIST using the same configuration as in the paper. The code can be referred :githublink:`here <examples/model_compress/pruning/lottery_torch_mnist_fc.py>`. In this experiment, we prune 10 times, for each pruning we train the pruned model for 50 epochs.
-.. image:: ../../img/lottery_ticket_mnist_fc.png
-   :target: ../../img/lottery_ticket_mnist_fc.png
-   :alt: 
-The above figure shows the result of the fully connected network. ``round0-sparsity-0.0`` is the performance without pruning. Consistent with the paper, pruning around 80% also obtain similar performance compared to non-pruning, and converges a little faster. If pruning too much, e.g., larger than 94%, the accuracy becomes lower and convergence becomes a little slower. A little different from the paper, the trend of the data in the paper is relatively more clear.
-Sensitivity Pruner
------------------
-For each round, SensitivityPruner prunes the model based on the sensitivity to the accuracy of each layer until meeting the final configured sparsity of the whole model:
-.. code-block:: bash
-       1. Analyze the sensitivity of each layer in the current state of the model.
-       2. Prune each layer according to the sensitivity.
-For more details, please refer to `Learning both Weights and Connections for Efficient Neural Networks  <https://arxiv.org/abs/1506.02626>`__.
-Usage
-^^^^^
-PyTorch code
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import SensitivityPruner
-   config_list = [{
-           'sparsity': 0.5,
-           'op_types': ['Conv2d']
-       }]
-   pruner = SensitivityPruner(model, config_list, finetuner=fine_tuner, evaluator=evaluator)
-   # eval_args and finetune_args are the parameters passed to the evaluator and finetuner respectively
-   pruner.compress(eval_args=[model], finetune_args=[model])
-User configuration for Sensitivity Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.SensitivityPruner
-Transformer Head Pruner
-----------------------
-Transformer Head Pruner is a tool designed for pruning attention heads from the models belonging to the `Transformer family <https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf>`__. The following image from `Efficient Transformers: A Survey <https://arxiv.org/pdf/2009.06732.pdf>`__ gives a good overview the general structure of the Transformer.
-.. image:: ../../img/transformer_structure.png
-   :target: ../../img/transformer_structure.png
-   :alt: 
-Typically, each attention layer in the Transformer models consists of four weights: three projection matrices for query, key, value, and an output projection matrix. The outputs of the former three matrices contains the projected results for all heads. Normally, the results are then reshaped so that each head performs that attention computation independently. The final results are concatenated back before fed into the output projection. Therefore, when an attention head is pruned, the same weights corresponding to that heads in the three projection matrices are pruned. Also, the weights in the output projection corresponding to the head's output are pruned. In our implementation, we calculate and apply masks to the four matrices together.
-Note: currently, the pruner can only handle models with projection weights written as separate ``Linear`` modules, i.e., it expects four ``Linear`` modules corresponding to query, key, value, and an output projections. Therefore, in the ``config_list``, you should either write ``['Linear']`` for the ``op_types`` field, or write names corresponding to ``Linear`` modules for the ``op_names`` field. For instance, the `Huggingface transformers <https://huggingface.co/transformers/index.html>`_ are supported, but ``torch.nn.Transformer`` is not.
-The pruner implements the following algorithm:
-.. code-block:: bash
-    Repeat for each pruning iteration (1 for one-shot pruning):
-       1. Calculate importance scores for each head in each specified layer using a specific criterion.
-       2. Sort heads locally or globally, and prune out some heads with lowest scores. The number of pruned heads is determined according to the sparsity specified in the config.
-       3. If the specified pruning iteration is larger than 1 (iterative pruning), finetune the model for a while before the next pruning iteration.
-Currently, the following head sorting criteria are supported:
-    * "l1_weight": rank heads by the L1-norm of weights of the query, key, and value projection matrices.
-    * "l2_weight": rank heads by the L2-norm of weights of the query, key, and value projection matrices.
-    * "l1_activation": rank heads by the L1-norm of their attention computation output.
-    * "l2_activation": rank heads by the L2-norm of their attention computation output.
-    * "taylorfo": rank heads by l1 norm of the output of attention computation * gradient for this output. Check more details in `this paper <https://arxiv.org/abs/1905.10650>`__ and `this one <https://arxiv.org/abs/1611.06440>`__.
-We support local sorting (i.e., sorting heads within a layer) and global sorting (sorting all heads together), and you can control by setting the ``global_sort`` parameter. Note that if ``global_sort=True`` is passed, all weights must have the same sparsity in the config list. However, this does not mean that each layer will be prune to the same sparsity as specified. This sparsity value will be interpreted as a global sparsity, and each layer is likely to have different sparsity after pruning by global sort. As a reminder, we found that if global sorting is used, it is usually helpful to use an iterative pruning scheme, interleaving pruning with intermediate finetuning, since global sorting often results in non-uniform sparsity distributions, which makes the model more susceptible to forgetting.
-In our implementation, we support two ways to group the four weights in the same layer together. You can either pass a nested list containing the names of these modules as the pruner's initialization parameters (usage below), or simply pass a dummy input instead and the pruner will run ``torch.jit.trace`` to group the weights (experimental feature). However, if you would like to assign different sparsity to each layer, you can only use the first option, i.e., passing names of the weights to the pruner (see usage below). Also, note that we require the weights belonging to the same layer to have the same sparsity.
-Usage
-^^^^^
-Suppose we want to prune a BERT with Huggingface implementation, which has the following architecture (obtained by calling ``print(model)``). Note that we only show the first layer of the repeated layers in the encoder's ``ModuleList layer``. 
-.. image:: ../../img/huggingface_bert_architecture.png
-   :target: ../../img/huggingface_bert_architecture.png
-   :alt: 
-**Usage Example: one-shot pruning, assigning sparsity 0.5 to the first six layers and sparsity 0.25 to the last six layers (PyTorch code)**. Note that
-* Here we specify ``op_names`` in the config list to assign different sparsity to different layers.
-* Meanwhile, we pass ``attention_name_groups`` to the pruner so that the pruner may group together the weights belonging to the same attention layer.
-* Since in this example we want to do one-shot pruning, the ``num_iterations`` parameter is set to 1, and the parameter ``epochs_per_iteration`` is ignored. If you would like to do iterative pruning instead, you can set the ``num_iterations`` parameter to the number of pruning iterations, and the ``epochs_per_iteration`` parameter to the number of finetuning epochs between two iterations.
-* The arguments ``trainer`` and ``optimizer`` are only used when we want to do iterative pruning, or the ranking criterion is ``taylorfo``. Here these two parameters are ignored by the pruner.
-* The argument ``forward_runner`` is only used when the ranking criterion is ``l1_activation`` or ``l2_activation``. Here this parameter is ignored by the pruner.
-.. code-block:: python
-   from nni.algorithms.compression.pytorch.pruning import TransformerHeadPruner
-   attention_name_groups = list(zip(["encoder.layer.{}.attention.self.query".format(i) for i in range(12)],
-                                    ["encoder.layer.{}.attention.self.key".format(i) for i in range(12)],
-                                    ["encoder.layer.{}.attention.self.value".format(i) for i in range(12)],
-                                    ["encoder.layer.{}.attention.output.dense".format(i) for i in range(12)]))
-   kwargs = {"ranking_criterion": "l1_weight",
-             "global_sort": False,
-             "num_iterations": 1,
-             "epochs_per_iteration": 1,    # this is ignored when num_iterations = 1
-             "head_hidden_dim": 64,
-             "attention_name_groups": attention_name_groups,
-             "trainer": trainer,
-             "optimizer": optimizer,
-             "forward_runner": forward_runner
-             }
-   config_list = [{
-	"sparsity": 0.5,
-        "op_types": ["Linear"],
-        "op_names": [x for layer in attention_name_groups[:6] for x in layer]      # first six layers
-   }, {
-	"sparsity": 0.25,
-	"op_types": ["Linear"],
-	"op_names": [x for layer in attention_name_groups[6:] for x in layer]      # last six layers
-   }]
-   pruner = TransformerHeadPruner(model, config_list, **kwargs)
-   pruner.compress()
-In addition to this usage guide, we provide a more detailed example of pruning BERT (Huggingface implementation) for transfer learning on the tasks from the `GLUE benchmark <https://gluebenchmark.com/>`_. Please find it in this :githublink:`page <examples/model_compress/pruning/transformers>`. To run the example, first make sure that you install the package ``transformers`` and ``datasets``. Then, you may start by running the following command:
-.. code-block:: bash
-	./run.sh gpu_id glue_task
-By default, the code will download a pretrained BERT language model, and then finetune for several epochs on the downstream GLUE task. Then, the ``TransformerHeadPruner`` will be used to prune out heads from each layer by a certain criterion (by default, the code lets the pruner uses magnitude ranking, and prunes out 50% of the heads in each layer in an one-shot manner). Finally, the pruned model will be finetuned in the downstream task for several epochs. You can check the details of pruning from the logs printed out by the example. You can also experiment with different pruning settings by changing the parameters in ``run.sh``, or directly changing the ``config_list`` in ``transformer_pruning.py``.
-User configuration for Transformer Head Pruner
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-**PyTorch**
-..  autoclass:: nni.algorithms.compression.pytorch.pruning.TransformerHeadPruner
--- a/docs/source/reference_zh.rst
+++ b/docs/source/reference_zh.rst
-.. 317504c3009932f8a566616e85a9700f
+.. e8dca0b3551823aef1648bcef1745028
 :orphan: