Issues 共 2585
KvCacheManagerV2 segfault on ~256k prompts with Nemotron-3-Ultra when max_seq_len=1M and avg_seq_len unset
#17926 · dsingal0 · 9 小时前
[Bug]: DGX Spark playbook supported models fail in trtllm-serve: Nemotron-3 Nano Omni and GPT-OSS rejected as unsupported due to MODEL_MAP registry mismatch
#15941 · evanrossd · 2026-07-05
[Bug]: Clamp very small non-zero temperature values to avoid numerical instability
#15715 · chfeng-cs · 2026-06-29
[Bug]: Gemma4 multimodal serving — two startup crashes (vision tp_size>1, xgrammar guided decoding)
#15613 · Thachnh · 2026-06-25
[AutoDeploy] Re-enable SSM replay for Nemotron-Super MTP (replay kernel illegal memory access at CUDA-graph capture on Blackwell)
#15565 · govind-ramnarayan · 2026-06-24
[Parity with vLLM, SGLang, ATOM]: Public Nightly NGC docker images too
#15328 · functionstackx · 2026-06-12
[Bug]: Qwen3.5-397B-A17B (bf16) OOMs host RAM on the PyTorch backend; AutoDeploy loads fine on the same node
#15296 · KyleShao1016 · 2026-06-12
[Feature Request] AutoDeploy: enable DeepSeek R1 MTP
#15277 · govind-ramnarayan · 2026-06-11
[Bug]: Guided decoding with Kimi-K2.5
#15179 · chungen04 · 2026-06-09
[Feature]: Investigate the unstable perf of llama 3.1 8B fp8 in CI test_sanity_perf.py
#15165 · MrGeva · 2026-06-09