Issues 共 1849
[Bug] requirements for the version of torch and triton
#4567 · zhangyexun · 2026-05-06
[Feature] TurboMind backend missing VLM support for DeepSeek-VL2, Llama4, Gemma3, Qwen3-VL
#4553 · ZhijunLStudio · 2026-04-24
[Feature] support DFlash: Block Diffusion for Flash Speculative Decoding
#4530 · hicofeng · 2026-04-15
[Feature] Support MiniMax-M2.7 in TurboMind engine
#4527 · bltcn · 2026-04-15
[Bug] qwen3.5推理TileLang依赖报错
#4512 · Zasa-Y · 2026-04-09
KV cache compression for longer context support
#4507 · jagmarques · 2026-04-07
Add TurboQuant Support for KV Cache Quantization
#4499 · janakg · 2026-04-06
[Bug] 使用lmdeploy在5090上cuda版本12.8,cuda toolkit同12.8
#4491 · lmingze · 2026-04-03
[Bug] lmdeploy run qwen3.5-122b-a10b-awq , which transformers version?
#4484 · wangchaoeric87 · 2026-04-01
[Feature] 现在可以支持Qwen3.5 4bit 量化吗?
#4464 · zhfeng1 · 2026-03-25