⬢github Python · 4.7K ★ +46 since we first saw it · pushed 1 h ago · Apache-2.0
NVIDIA/Model-Optimizer
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
NVIDIA Model Optimizer (ModelOpt) is a Python library that compresses deep learning models using techniques like quantization, pruning, distillation, sparsity, neural architecture search, and speculative decoding. It takes Hugging Face, PyTorch, or ONNX models, applies optimization via Python APIs, and exports optimized checkpoints ready for deployment in inference frameworks like TensorRT-LLM, TensorRT, vLLM, and SGLang for faster inference.
Why now: Recent NVFP4 quantization news and blogs, including Nemotron 3 Ultra 550B NVFP4 checkpoints and W4A4 tutorials, plus new AutoQuantize and quantization-aware distillation posts.
Who it is for: ML engineers deploying large models on NVIDIA hardware who want smaller, faster inference without rewriting their training pipeline.
Stars over our 15 snapshots: 4.7K to 4.7K, since 1 h ago.
Where people talked about it
- ⬢github NVIDIA/Model-Optimizer 3 min ago
API: https://socialmediatrends-api.osmike.com/v1/repos/NVIDIA/Model-Optimizer