Issues 共 396
[Torchtitan][Pytorch][ROCm 7.14] Qwen3 MoE TP+EP training produces extremely large(nan) gradient norms
#4205 · vidushi8 · 23 小时前
[rl] vLLM torch.compile (VLLM_COMPILE) on the generator fails: inductor "out variant should have at least one out arg"
#3669 · yichuan-w · 2026-06-15
FlexAttention `BlockMask` breaks Pipeline Parallelism with split-backward (zero-bubble / multi) schedules
#3584 · tianyu-l · 2026-06-09
inexistent etp flag in the experiments/ folder
#3566 · JINO-ROHIT · 2026-06-07
CI: 8 GPU Integration Test on H100 ROCm leg red on main (blocked on pytorch/test-infra#8152)
#3552 · rishisinhanj · 2026-06-05
[RL][Flex] batch invariant Flex Attention doesn't work with general compile
#3497 · liangel-02 · 2026-06-03
HybridEP DispatchHandle not working with torch.compile
#3439 · chelsea0x3b · 2026-05-28
[gpt_oss] Avoid materialising `(x_linear + 1)` intermediate in `swiglu` to reduce activation memory
#3437 · janbernloehr · 2026-05-27
DeepEP: CUDA stream race between masked_fill and get_dispatch_layout causes grouped_mm assertion failure
#3413 · acisseJZhong · 2026-05-21
torch.compile fails with MoE + EP: PendingUnbackedSymbolNotFound on DTensor .view(-1, dim)
#3409 · acisseJZhong · 2026-05-20