global_automodel#
Classes
Functions
Translate Puzzletron KD settings into the current NeMo AutoModel recipe schema. |
|
- class GlobalKDConfig#
Bases:
object- __init__(*, teacher_dir, student_dir, output_dir, descriptor=None, teacher_descriptor=None, student_descriptor=None, force_hf=True, teacher_force_hf=None, student_force_hf=None, teacher_model_kwargs=<factory>, student_model_kwargs=<factory>, domain='auto', trust_remote_code=True, torch_dtype='bf16', attn_implementation=None, tp=1, pp=1, ep=1, cp=1, dp=1, sequence_parallel=False, activation_checkpointing=False, pp_schedule='1f1b', save_consolidated=False, checkpoint_format='auto', main_ce_weight=1.0, mtp_ce_weight=0.0, main_kd_weight=1.0, mtp_kd_weight=0.0, main_kd=<factory>, mtp_kd=<factory>, ce_weight=None, kd_weight=None, hidden_kd_weight=0.0, temperature=None, global_batch_size=128, local_batch_size=1, max_steps=1, checkpoint_every_steps=None, packed_sequence_size=0, lr=1e-05, weight_decay=0.0, seed=1111, dataset_name='rajpurkar/squad', dataset_split='train', validation_enabled=True, validation_split='validation', resume=True, metadata=<factory>, data=<factory>, freeze_policy='vision_frozen')#
- Parameters:
teacher_dir (Path)
student_dir (Path)
output_dir (Path)
descriptor (str | None)
teacher_descriptor (str | None)
student_descriptor (str | None)
force_hf (bool)
teacher_force_hf (bool | None)
student_force_hf (bool | None)
teacher_model_kwargs (dict[str, Any])
student_model_kwargs (dict[str, Any])
domain (Literal['auto', 'llm', 'vlm'])
trust_remote_code (bool)
torch_dtype (str)
attn_implementation (str | None)
tp (int)
pp (int)
ep (int)
cp (int)
dp (int)
sequence_parallel (bool)
activation_checkpointing (bool | str)
pp_schedule (Literal['1f1b', 'interleaved1f1b'])
save_consolidated (bool | str)
checkpoint_format (Literal['auto', 'safetensors', 'torch_save'])
main_ce_weight (float)
mtp_ce_weight (float)
main_kd_weight (float)
mtp_kd_weight (float)
main_kd (KDLossTermConfig)
mtp_kd (KDLossTermConfig)
ce_weight (float | None)
kd_weight (float | None)
hidden_kd_weight (float)
temperature (float | None)
global_batch_size (int)
local_batch_size (int)
max_steps (int)
checkpoint_every_steps (int | None)
packed_sequence_size (int)
lr (float)
weight_decay (float)
seed (int)
dataset_name (str)
dataset_split (str)
validation_enabled (bool)
validation_split (str)
resume (bool)
metadata (dict[str, Any])
data (dict[str, Any])
freeze_policy (Literal['vision_frozen', 'projector_and_language', 'train_all'])
- Return type:
None
- activation_checkpointing: bool | str = False#
- attn_implementation: str | None = None#
- ce_weight: float | None = None#
- checkpoint_every_steps: int | None = None#
- checkpoint_format: Literal['auto', 'safetensors', 'torch_save'] = 'auto'#
- cp: int = 1#
- data: dict[str, Any]#
- dataset_name: str = 'rajpurkar/squad'#
- dataset_split: str = 'train'#
- descriptor: str | None = None#
- domain: Literal['auto', 'llm', 'vlm'] = 'auto'#
- dp: int = 1#
- ep: int = 1#
- force_hf: bool = True#
- freeze_policy: Literal['vision_frozen', 'projector_and_language', 'train_all'] = 'vision_frozen'#
- global_batch_size: int = 128#
- property identity: str#
- kd_weight: float | None = None#
- local_batch_size: int = 1#
- lr: float = 1e-05#
- main_ce_weight: float = 1.0#
- main_kd: KDLossTermConfig#
- main_kd_weight: float = 1.0#
- max_steps: int = 1#
- metadata: dict[str, Any]#
- mtp_ce_weight: float = 0.0#
- mtp_kd: KDLossTermConfig#
- mtp_kd_weight: float = 0.0#
- property needs_teacher: bool#
- property objective_weights: dict[str, float]#
- output_dir: Path#
- packed_sequence_size: int = 0#
- pp: int = 1#
- pp_schedule: Literal['1f1b', 'interleaved1f1b'] = '1f1b'#
- property resolved_student_descriptor: str | None#
- property resolved_student_force_hf: bool#
- property resolved_teacher_descriptor: str | None#
- property resolved_teacher_force_hf: bool#
- resume: bool = True#
- save_consolidated: bool | str = False#
- seed: int = 1111#
- sequence_parallel: bool = False#
- student_descriptor: str | None = None#
- student_dir: Path#
- student_force_hf: bool | None = None#
- student_model_kwargs: dict[str, Any]#
- teacher_descriptor: str | None = None#
- teacher_dir: Path#
- teacher_force_hf: bool | None = None#
- teacher_model_kwargs: dict[str, Any]#
- temperature: float | None = None#
- to_dict()#
- Return type:
dict[str, Any]
- torch_dtype: str = 'bf16'#
- tp: int = 1#
- trust_remote_code: bool = True#
- validation_enabled: bool = True#
- validation_split: str = 'validation'#
- weight_decay: float = 0.0#
- class GlobalKDResult#
Bases:
object- __init__(*, kd_id, output_dir, metrics=<factory>)#
- Parameters:
kd_id (str)
output_dir (Path)
metrics (dict[str, Any])
- Return type:
None
- kd_id: str#
- metrics: dict[str, Any]#
- output_dir: Path#
- to_dict()#
- Return type:
dict[str, Any]
- class KDLossTermConfig#
Bases:
object- __init__(*, metric='kld', temperature=1.0, chunk_size=0)#
- Parameters:
metric (Literal['kld', 'tvd'])
temperature (float)
chunk_size (int)
- Return type:
None
- chunk_size: int = 0#
- metric: Literal['kld', 'tvd'] = 'kld'#
- temperature: float = 1.0#
- build_automodel_global_kd_recipe(kd_config)#
Translate Puzzletron KD settings into the current NeMo AutoModel recipe schema.
- Parameters:
kd_config (GlobalKDConfig)
- Return type:
dict[str, Any]
- build_global_kd_config(config)#
- Parameters:
config (dict[str, Any])
- Return type:
- run_automodel_global_kd(kd_config)#
- Parameters:
kd_config (GlobalKDConfig)
- Return type:
dict[str, Any]
- run_global_kd(kd_config, recipe_runner=None)#
- Parameters:
kd_config (GlobalKDConfig)
- Return type: