lora#
Mergeable low-rank adaptation of fake-quantized linear weights.
Classes
Configuration for adapters applied to combined weights before fake quantization. |
Functions
Freeze a fake-quantized backbone and add zero-initialized, mergeable adapters. |
|
Merge adapters into floating-point weights, retaining quantizers for deployment export. |
- class QuantLoRAConfig#
Bases:
ModeloptBaseConfigConfiguration for adapters applied to combined weights before fake quantization.
- alpha: float#
- model_config = {'extra': 'forbid', 'validate_assignment': True}#
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- rank: int#
- target_modules: list[str]#
- enable_quant_lora(model, config=None)#
Freeze a fake-quantized backbone and add zero-initialized, mergeable adapters.
The forward quantizes
W + (alpha / rank) * B @ A. Save and restore using ModelOpt checkpoint APIs to reconstruct adapters before loading tensor/optimizer state.- Parameters:
model (Module)
config (dict | QuantLoRAConfig | None)
- Return type:
Module
- merge_quant_lora(model)#
Merge adapters into floating-point weights, retaining quantizers for deployment export.
- Parameters:
model (Module)
- Return type:
Module