quant_linear#
Quantized Linear.
Classes
Quantized version of nn.Linear. |
|
Quantized version of nn.Linear with real quantization. |
|
Base class for quantized linear modules with SVDQuant. |
- Linear#
alias of
QuantLinear
- class QuantLinear#
Bases:
_LegacyQuantLinearConvBaseMixin,LinearQuantized version of nn.Linear.
- default_quant_desc_weight = QuantizerAttributeConfig(enable=True, num_bits=8, effective_bits=None, axis=0, fake_quant=True, unsigned=False, narrow_range=False, learn_amax=False, type='static', block_sizes=None, bias=None, trt_high_precision_dtype='Float', calibrator='max', rotate=False, pass_through_bwd=True, backend=None, backend_extra_args=None, use_constant_amax=False, constant_amax=None)#
- class RealQuantLinear#
Bases:
QuantModuleQuantized version of nn.Linear with real quantization.
- allow_real_quant_gemm = True#
- forward(input, *args, **kwargs)#
RealQuant layer forward function.
- has_real_quant_gemm_impl(input, *args, **kwargs)#
Get the real quant GEMM implementation base on input arguments.
- Return type:
bool
- list_of_scale_tensors = ['_scale', '_double_scale', '_scale_zeros']#
- class SVDQuantLinear#
Bases:
QuantLinearConvBaseBase class for quantized linear modules with SVDQuant.
- fold_weight(keep_attrs=False)#
Fold the weight for faster eval.
- Parameters:
keep_attrs (bool)
- forward(input, *args, **kwargs)#
SVDQuant layer forward function.