ggml#

GGML-compatible block quantization formats.

Classes

GGMLFormat

Everything backend dispatch and export need to know about one GGML format.

IQFormat

Functions

dequantize_iq1_m

Decode GGML-compatible IQ1_M payload bytes.

iq1_m_fake_quant

TensorQuantizer backend for this format, with pass-through backward.

iq1_m_grid

Return the canonical IQ1_M ternary grid as float32.

quantize_iq1_m

Pack a floating-point weight into GGML-compatible IQ1_M blocks.

dequantize_iq1_s

Decode GGML-compatible IQ1_S payload bytes.

iq1_s_fake_quant

TensorQuantizer backend for this format, with pass-through backward.

iq1_s_grid

Return the canonical IQ1_S ternary grid as float32.

quantize_iq1_s

Pack a floating-point weight into GGML-compatible IQ1_S blocks.

dequantize_iq2_s

Decode GGML-compatible IQ2_S payload bytes.

iq2_s_fake_quant

TensorQuantizer backend for this format, with pass-through backward.

iq2_s_grid

Return the canonical IQ2_S magnitude grid as float32.

quantize_iq2_s

Pack a floating-point weight into GGML-compatible IQ2_S blocks.

dequantize_iq2_xs

Decode GGML-compatible IQ2_XS payload bytes.

iq2_xs_fake_quant

TensorQuantizer backend for this format, with pass-through backward.

iq2_xs_grid

Return the canonical IQ2_XS magnitude grid as float32.

quantize_iq2_xs

Pack a floating-point weight into GGML-compatible IQ2_XS blocks.

dequantize_iq2_xxs

Decode GGML-compatible IQ2_XXS payload bytes.

iq2_xxs_fake_quant

TensorQuantizer backend for this format, with pass-through backward.

iq2_xxs_grid

Return the canonical IQ2_XXS magnitude grid as float32.

quantize_iq2_xxs

Pack a floating-point weight into GGML-compatible IQ2_XXS blocks.

dequantize_q8_0

Decode GGML-compatible Q8_0 payload bytes.

q8_0_fake_quant

TensorQuantizer backend for this format, with pass-through backward.

quantize_q8_0

Pack a floating-point weight into GGML-compatible Q8_0 blocks.

class GGMLFormat#

Bases: object

Everything backend dispatch and export need to know about one GGML format.

Each format module declares one of these beside its encoder and decoder, and GGML_FORMAT_REGISTRY lists them. The per-format pieces – codebook, search, payload layout – stay in the format’s module; what lives here is the part every format does the same way.

__init__(name, block_size, block_bytes, quantize, dequantize, block_chunk_size, decode_chunk_size)#
Parameters:
  • name (str)

  • block_size (int)

  • block_bytes (int)

  • quantize (Callable[[...], tuple[Tensor, Tensor]])

  • dequantize (Callable[[...], Tensor])

  • block_chunk_size (int)

  • decode_chunk_size (int)

Return type:

None

block_bytes: int#
block_chunk_size: int#
block_size: int#
decode_chunk_size: int#
dequantize: Callable[[...], Tensor]#
property effective_bits: float#

Packed storage cost per weight.

fake_quant(inputs, quantizer, *, block_chunk_size=None, decode_chunk_size=None)#

TensorQuantizer backend for this format, with pass-through backward.

Parameters:
  • inputs (Tensor)

  • block_chunk_size (int | None)

  • decode_chunk_size (int | None)

Return type:

Tensor

name: str#
pack(weight, quantizer=None)#

The packed payload of weight, reusing the one fake quant cached if it matches.

Export calls this. A model that ran a forward since quantization already holds each weight’s payload, and reusing it both skips a second search and exports exactly the bytes the evaluated model decoded. Any other tensor is packed afresh, or takes the payload pinned to quantizer if it is that payload’s decoding.

Parameters:

weight (Tensor)

Return type:

Tensor

quantize: Callable[[...], tuple[Tensor, Tensor]]#
IQFormat#

alias of GGMLFormat

dequantize_iq1_m(packed_weights, weight_shape, *, dtype=torch.bfloat16, block_chunk_size=4096)#

Decode GGML-compatible IQ1_M payload bytes.

Parameters:
  • packed_weights (Tensor)

  • weight_shape (Tensor)

  • dtype (dtype)

  • block_chunk_size (int)

Return type:

Tensor

dequantize_iq1_s(packed_weights, weight_shape, *, dtype=torch.bfloat16, block_chunk_size=4096)#

Decode GGML-compatible IQ1_S payload bytes.

Parameters:
  • packed_weights (Tensor)

  • weight_shape (Tensor)

  • dtype (dtype)

  • block_chunk_size (int)

Return type:

Tensor

dequantize_iq2_s(packed_weights, weight_shape, *, dtype=torch.bfloat16, block_chunk_size=4096)#

Decode GGML-compatible IQ2_S payload bytes.

Parameters:
  • packed_weights (Tensor)

  • weight_shape (Tensor)

  • dtype (dtype)

  • block_chunk_size (int)

Return type:

Tensor

dequantize_iq2_xs(packed_weights, weight_shape, *, dtype=torch.bfloat16, block_chunk_size=4096)#

Decode GGML-compatible IQ2_XS payload bytes.

Parameters:
  • packed_weights (Tensor)

  • weight_shape (Tensor)

  • dtype (dtype)

  • block_chunk_size (int)

Return type:

Tensor

dequantize_iq2_xxs(packed_weights, weight_shape, *, dtype=torch.bfloat16, block_chunk_size=4096)#

Decode GGML-compatible IQ2_XXS payload bytes.

Parameters:
  • packed_weights (Tensor)

  • weight_shape (Tensor)

  • dtype (dtype)

  • block_chunk_size (int)

Return type:

Tensor

dequantize_q8_0(packed_weights, weight_shape, *, dtype=torch.bfloat16, block_chunk_size=16384)#

Decode GGML-compatible Q8_0 payload bytes.

Parameters:
  • packed_weights (Tensor)

  • weight_shape (Tensor)

  • dtype (dtype)

  • block_chunk_size (int)

Return type:

Tensor

iq1_m_fake_quant(inputs, quantizer, *, block_chunk_size=None, decode_chunk_size=None)#

TensorQuantizer backend for this format, with pass-through backward.

Parameters:
  • inputs (Tensor)

  • block_chunk_size (int | None)

  • decode_chunk_size (int | None)

Return type:

Tensor

iq1_m_grid(device=None)#

Return the canonical IQ1_M ternary grid as float32.

IQ1_M indexes the same 2048-entry table as IQ1_S; this alias exists so every format exposes a grid accessor under its own name.

Parameters:

device (device | str | None)

Return type:

Tensor

iq1_s_fake_quant(inputs, quantizer, *, block_chunk_size=None, decode_chunk_size=None)#

TensorQuantizer backend for this format, with pass-through backward.

Parameters:
  • inputs (Tensor)

  • block_chunk_size (int | None)

  • decode_chunk_size (int | None)

Return type:

Tensor

iq1_s_grid(device=None)#

Return the canonical IQ1_S ternary grid as float32.

Parameters:

device (device | str | None)

Return type:

Tensor

iq2_s_fake_quant(inputs, quantizer, *, block_chunk_size=None, decode_chunk_size=None)#

TensorQuantizer backend for this format, with pass-through backward.

Parameters:
  • inputs (Tensor)

  • block_chunk_size (int | None)

  • decode_chunk_size (int | None)

Return type:

Tensor

iq2_s_grid(device=None)#

Return the canonical IQ2_S magnitude grid as float32.

Parameters:

device (device | str | None)

Return type:

Tensor

iq2_xs_fake_quant(inputs, quantizer, *, block_chunk_size=None, decode_chunk_size=None)#

TensorQuantizer backend for this format, with pass-through backward.

Parameters:
  • inputs (Tensor)

  • block_chunk_size (int | None)

  • decode_chunk_size (int | None)

Return type:

Tensor

iq2_xs_grid(device=None)#

Return the canonical IQ2_XS magnitude grid as float32.

Parameters:

device (device | str | None)

Return type:

Tensor

iq2_xxs_fake_quant(inputs, quantizer, *, block_chunk_size=None, decode_chunk_size=None)#

TensorQuantizer backend for this format, with pass-through backward.

Parameters:
  • inputs (Tensor)

  • block_chunk_size (int | None)

  • decode_chunk_size (int | None)

Return type:

Tensor

iq2_xxs_grid(device=None)#

Return the canonical IQ2_XXS magnitude grid as float32.

Parameters:

device (device | str | None)

Return type:

Tensor

q8_0_fake_quant(inputs, quantizer, *, block_chunk_size=None, decode_chunk_size=None)#

TensorQuantizer backend for this format, with pass-through backward.

Parameters:
  • inputs (Tensor)

  • block_chunk_size (int | None)

  • decode_chunk_size (int | None)

Return type:

Tensor

quantize_iq1_m(weight, *, block_chunk_size=1024)#

Pack a floating-point weight into GGML-compatible IQ1_M blocks.

Returned shapes are [*weight.shape[:-1], weight.shape[-1] // 256, 56] and [weight.ndim].

Parameters:
  • weight (Tensor)

  • block_chunk_size (int)

Return type:

tuple[Tensor, Tensor]

quantize_iq1_s(weight, *, block_chunk_size=1024)#

Pack a floating-point weight into GGML-compatible IQ1_S blocks.

Returned shapes are [*weight.shape[:-1], weight.shape[-1] // 256, 50] and [weight.ndim]. The packed payload remains on the weight’s device; the logical-shape metadata is kept on CPU. Non-finite input elements are treated as zero during packing.

Parameters:
  • weight (Tensor)

  • block_chunk_size (int)

Return type:

tuple[Tensor, Tensor]

quantize_iq2_s(weight, *, block_chunk_size=128)#

Pack a floating-point weight into GGML-compatible IQ2_S blocks.

Returned shapes are [*weight.shape[:-1], weight.shape[-1] // 256, 82] and [weight.ndim].

Parameters:
  • weight (Tensor)

  • block_chunk_size (int)

Return type:

tuple[Tensor, Tensor]

quantize_iq2_xs(weight, *, block_chunk_size=256)#

Pack a floating-point weight into GGML-compatible IQ2_XS blocks.

Returned shapes are [*weight.shape[:-1], weight.shape[-1] // 256, 74] and [weight.ndim]. The packed payload remains on the weight’s device; the logical-shape metadata is kept on CPU. Non-finite input elements are treated as zero during packing.

Parameters:
  • weight (Tensor)

  • block_chunk_size (int)

Return type:

tuple[Tensor, Tensor]

quantize_iq2_xxs(weight, *, block_chunk_size=512)#

Pack a floating-point weight into GGML-compatible IQ2_XXS blocks.

Returned shapes are [*weight.shape[:-1], weight.shape[-1] // 256, 66] and [weight.ndim]. The packed payload remains on the weight’s device; the logical-shape metadata is kept on CPU. Non-finite input elements are treated as zero during packing.

Parameters:
  • weight (Tensor)

  • block_chunk_size (int)

Return type:

tuple[Tensor, Tensor]

quantize_q8_0(weight, *, block_chunk_size=4096)#

Pack a floating-point weight into GGML-compatible Q8_0 blocks.

Parameters:
  • weight (Tensor)

  • block_chunk_size (int)

Return type:

tuple[Tensor, Tensor]