lora#

Mergeable low-rank adaptation of fake-quantized linear weights.

Classes

QuantLoRAConfig

Configuration for adapters applied to combined weights before fake quantization.

Functions

enable_quant_lora

Freeze a fake-quantized backbone and add zero-initialized, mergeable adapters.

merge_quant_lora

Merge adapters into floating-point weights, retaining quantizers for deployment export.

class QuantLoRAConfig#

Bases: ModeloptBaseConfig

Configuration for adapters applied to combined weights before fake quantization.

alpha: float#
model_config = {'extra': 'forbid', 'validate_assignment': True}#

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

rank: int#
target_modules: list[str]#
enable_quant_lora(model, config=None)#

Freeze a fake-quantized backbone and add zero-initialized, mergeable adapters.

The forward quantizes W + (alpha / rank) * B @ A. Save and restore using ModelOpt checkpoint APIs to reconstruct adapters before loading tensor/optimizer state.

Parameters:
Return type:

Module

merge_quant_lora(model)#

Merge adapters into floating-point weights, retaining quantizers for deployment export.

Parameters:

model (Module)

Return type:

Module