search
vLLM
Trends
- 1
NVIDIA/Model-Optimizer is an open-source Python library on GitHub that collects state-of-the-art model optimization techniques, including quantization, distillation, pruning, neural architecture search and speculative decoding. It compresses deep learning models so they run efficiently in deployment frameworks such as TensorRT-LLM, TensorRT and vLLM, improving inference speed. It is trending on GitHub's rankings with modest engagement, and the posts shown only describe the project itself, so there is no evidence of a specific event driving attention.
Repos
- Contrastive-LM/CLM
- nokia-applied-research/AnyJev Turn any LLM into a Jev-style decision model: typed decisions, real probabilities, no training. (continue updating, welc
- QwenLM/Qwen-Image-2.1 Qwen's most powerful open-source image generation model
- githubnext/localjev
- NVIDIA/Model-Optimizer A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture se