Issues 共 506
Runtime check for the FlashAttention version rejects local version identifiers when it shouldn't
#3334 · lupreCSC · 11 天前
[BUG] NVFP4BlockScaling forward pass fails on RTX 5080 (sm_120) when K is a multiple of 128; sm_100/B100 TMEM+UMMA hardcoded in row_cast_col_hadamard_transform_cast_fusion.cu
#2956 · osubotin · 2026-05-03
[BUG]TensorRT engine export fails with UNSUPPORTED_NODE errors after TE FP8 training
#2915 · charriry · 2026-04-22
[Bug] moe_permute CUDA kernel: int32 overflow and incorrect -1 sentinel handling
#2908 · jing-4369 · 2026-04-21
FP8 weight caching shows no speedup (sometimes slowdown) and causes training curve differences under Float8BlockScaling
#2852 · Mnb66 · 2026-04-08
Request for batched general_gemm() (or FP8-aware torch.bmm) for non-Linear GEMM workloads
#2846 · jomitchellnv · 2026-04-07
[Core] MXFP8 grouped quantization kernel crashes with groups of size zero
#2779 · jberchtold-nvidia · 2026-03-18
undefined symbol: _ZN3c104cuda29c10_cuda_check_implementationEiPKcS2_ib
#2771 · ZMY-0309 · 2026-03-17
Grouped Bias/Dbias Kernel Support After Grouped GEMM
#2766 · vthumbe1503 · 2026-03-16
[Question] Expected behavior for blockwise FP8? Hybrid E4M3/E5M2 format & eval metrics outperforming BF16
#2754 · Mnb66 · 2026-03-11