-
Notifications
You must be signed in to change notification settings - Fork 506
Pull requests: NVIDIA/Model-Optimizer
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Speed up compressed-tensors load-time matching (for Kimi models)
#1999
opened Jul 21, 2026 by
rohansjoshi
Contributor
Loading…
fix(export): prevent silent MLP linear collapse during HF export
#1997
opened Jul 20, 2026 by
Surya-5555
Loading…
feat(rocm): Add AMD ROCm/MI300X support — FP8 hipBLASLt, MIGraphX backend, AMD quantization configs
#1990
opened Jul 17, 2026 by
zhihuidu-amd
Loading…
[6425069][ONNX][Autocast] Fix autocast metadata propagation
#1983
opened Jul 16, 2026 by
gcunhase
Contributor
Loading…
launcher: bump TRT-LLM to 1.3.0rc20, pin vLLM to v0.22.0, fix max_seq…
#1982
opened Jul 16, 2026 by
noeyy-mino
Contributor
Loading…
[5726458] Add NVFP4 projection-output-quantizer recipe and HF embedding ONNX export example
#1981
opened Jul 16, 2026 by
ajrasane
Contributor
Loading…
Scripts and a skill to do per-layer benchmark using flashinfer
#1980
opened Jul 16, 2026 by
sychen52
Contributor
Loading…
Quality: Insecure subprocess usage in get_system_info.py
#1977
opened Jul 15, 2026 by
tomaioo
Loading…
Multi-GPU vLLM benchmarking for runtime stats
#1972
opened Jul 14, 2026 by
grzegorz-k-karch
Contributor
•
Draft
Previous Next
ProTip!
no:milestone will show everything without a milestone.