ITADN
flashinfer-ai/flashinfer

Issues 465

[Bug] JIT: bgmv_moe rebuilds from scratch in every process because shutil.copy resets the copied sources' mtime
#4782 · yufeiwu-nv · 1 天前
[Bug] flashinfer-cubin wheels missing on PyPI for >=0.6.14 — version check breaks every vLLM >= 0.27 on DGX Spark (GB10/aarch64)
#4781 · skbotoc1-web · 1 天前
[Feature]: yo sis can you create like extend the task scheduled kernal coverage for custom attention operations
#4777 · kds1123001 · 1 天前
[Bug][create-release] ARM64 cu129 flashinfer-jit-cache wheel exceeds the 2 GiB GitHub Release asset limit
#4519 · MaryWang-Nvidia · 15 天前
[Bug] Paged causal prefill returns finite LSE for fully masked rows when qo_len > kv_len
#4452 · anguyen8 · 18 天前
[Bug][v0.6.17][gb300]tests/gemm/test_groupwise_scaled_gemm_fp8.py:201: Mismatched elements: 124 / 8192 (1.5%)
#4396 · nvamyt · 21 天前
[bug][v0.6.17][rtx_5090] tests.moe.test_unified_moe_b12x.TestB12xUnifiedValidation.test_b12x_does_not_support_unfinalized_output TypeError: TestB12xUnifiedValidation._config() got an unexpected keyword argument 'finalize'
#4395 · nvamyt · 21 天前
[Bug] test_proxy_split_k_heuristic fails on RTX 5090 (170 SM)
#4184 · kangbintNV · 2026-07-27
SM103 serving hang caused by TRTLLM_GEN_BMM artifact regeneration in #3708 (batched_gemm-dd6d23e-721ae60, 0.6.14): FP4 batched-GEMM clusters stuck in mbarrier phase wait
#3971 · YAMY1234 · 2026-07-14
[Bug] cuDNN FP8 MOE grouped matmul silently computes only the first expert group with backend 9.18–9.20 (SM120)
#3792 · waynehacking8 · 2026-07-01