ReleaseNVIDIANVIDIApublished May 13, 2026seen 5d

NVIDIA/Model-Optimizer 0.44.0

NVIDIA/Model-Optimizer

Open original ↗

Captured source

source ↗
published May 13, 2026seen 5dcaptured 15hhttp 200method plain

ModelOpt 0.44.0 Release

Repository: NVIDIA/Model-Optimizer

Tag: 0.44.0

Published: 2026-05-13T20:37:10Z

Prerelease: no

Release notes:

New Features

  • Support full Transformer Engine spec for Minitron pruning (mcore_minitron). Now we no longer need to use custom ModelOpt spec. Note that this does not affect the usage of the pruning workflow but makes pruning slightly faster and may result in slightly different pruned model because of different kernel and numerics.
  • Add end-to-end tutorial for Minitron pruning + distillation + quantization + evaluation + vLLM deployment for Nemotron-Nano-9B-v2 → Pruned 7B along with data blend preparation steps (and ablation study). See examples/pruning/minitron/README.md for details.
  • Add Puzzletron - a new algorithm for heterogeneous pruning of LLM and VLM models. See examples/puzzletron/README.md for more details.
  • Added iterator interface using CalibrationDataReader in ONNX quantization workflow.
  • Add N:M sparse softmax support to the Triton flash attention kernel (modelopt.torch.kernels.common.attention.triton_fa). See examples/llm_sparsity/attention_sparsity/README.md for usage.
  • Add skip-softmax skipping to the Triton flash attention kernel (modelopt.torch.kernels.common.attention.triton_fa). See examples/llm_sparsity/attention_sparsity/README.md for usage.
  • Add Video Sparse Attention (VSA) method for video diffusion models (modelopt.torch.sparsity.attention_sparsity). VSA uses 3D block tiling with a two-branch architecture for attention speedup.
  • Enable PTQ workflow for the Step3.5-Flash MoE model with NVFP4 W4A4 + FP8 KV cache quantization. See modelopt_recipes/models/Step3.5-Flash/nvfp4-mlp-only.yaml for more details.
  • Add support for vLLM fakequant reload using ModelOpt state for HF models. See examples/vllm_serve/README.md for more details.
  • [Early Testing] Add Claude Code PTQ skill (.claude/skills/ptq/) for agent-assisted post-training quantization. The skill guides the agent through environment detection, model support checking, format selection, and execution via the launcher or manual SLURM/Docker/bare GPU paths. Includes handling for unlisted models with custom module patching. This feature is in early testing — use with caution.
  • [Early Testing] Polish Claude Code evaluation skill (.claude/skills/evaluation/) for agent-assisted LLM accuracy benchmarking via NeMo Evaluator Launcher. Adds two companion skills vendored verbatim from NVIDIA-NeMo/Evaluator: launching-evals (run/check/debug/analyze NEL evaluations) and accessing-mlflow (query MLflow runs, compare metrics, fetch artifacts). Re-sync at a pinned upstream SHA via .claude/scripts/sync-upstream-skills.sh. Also adds a shared skills/common/credentials.md covering HF / NGC / Docker token setup referenced by multiple skills. This feature is in early testing — use with caution.
  • Add performant layerwise calibration for large models that don't fit on GPU (e.g. DeepSeek-R1, Kimi-K2). See modelopt_recipes/general/ptq/nvfp4_experts_only-kv_fp8_layerwise.yaml for usage. Layerwise calibration also supports PTQ with intermediate progress saving — useful when long PTQ runs get hit with Slurm timeouts. See modelopt_recipes/general/ptq/nvfp4_default-kv_none-gptq.yaml for usage.
  • Add implicit GEMM CUDA kernel for Conv3D with fused NVFP4 fake quantization (modelopt.torch.quantization.src.conv). When NVFP4 quantization is applied to an nn.Conv3d layer via ModelOpt PTQ, the implicit GEMM path is used automatically instead of cuDNN. Uses BF16 WMMA tensor cores (SM80+) with FP32 accumulation and in-kernel FP4 (E2M1) activation quantization. Grouped convolution (groups > 1) falls back to the default cuDNN path. Inference only — training mode falls back to cuDNN with a warning.
  • Add FP8 MHA quantization support for vision transformers. Adds an attention-aware ONNX post-processing pass (scale Mul / K-transpose move before Q, Q→DQ insertion on softmax output) in FP8QuantExporter (modelopt.onnx.export.fp8_exporter.FP8QuantExporter), per-instance nested-attention-wrapper skipping in the HF plugin, and nn.LayerNorm registration in QuantModuleRegistry so BMM input quantizers and LayerNorm output quantizers defined in FP8_DEFAULT_CFG are honored end-to-end. See examples/torch_onnx/torch_quant_to_onnx.py for the general timm-model quantize→ONNX workflow.

Backward Breaking Changes

  • The quant_cfg field in quantization configs is now an ordered list of QuantizerCfgEntry dicts instead of a flat dictionary. Each entry specifies a quantizer_name wildcard, an optional parent_class filter, a cfg dict of quantizer attributes, and/or an enable flag. Entries are applied in list order with later entries overriding earlier ones. The old dict-based format is still accepted and automatically converted via normalize_quant_cfg_list(), but now emits a DeprecationWarning; new code should use the list format. All built-in configs (e.g. FP8_DEFAULT_CFG, INT4_AWQ_CFG, NVFP4_DEFAULT_CFG), examples, and YAML recipes have been updated. See the quant-cfg documentation for the new format reference and migration guide.
  • Deprecated Mllama (Llama 3.2 Vision) support in the llm_ptq and vlm_ptq examples. The model_type == "mllama" branches and MllamaImageProcessor usage have been removed from hf_ptq.py and example_utils.py. For image-text calibration of VLMs, use --calib_with_images with a supported VLM (see Nemotron VL section in examples/llm_ptq/README.md).

Bug Fixes

  • Fix Megatron…

Excerpt shown — open the source for the full document.

Notability

notability 4.0/10

Routine minor release of an optimization tool.