Issues 共 438
Data-dependent bidirectional mask (Gemma-4 vision-block) forces standard Attention over GQA — split prefill/decode decoder graphs?
#2204 · justinchuby · 2026-06-09
BFloat16 logits returned as garbage — Logits::Get() only casts Float16 to float32
#2202 · justinchuby · 2026-06-09
OGA export generating an invalid model when a lm_head unquantized model is exported
#2185 · uday610 · 2026-05-27
Copilot design of Pipeline-as-Config
#2114 · justinchuby · 2026-05-02
Feature Request: Support for Google Gemma 4 model family (PLE architecture, variable head dims, KV cache sharing)
#2062 · elbruno · 2026-04-03
cannot build Qwen/Qwen3-14B on machine with 64GB RAM
#2047 · xiaofeihan1 · 2026-03-26
[Java] `onnxruntime-genai-cuda.dll` missing from Java JAR and extraction logic in CUDA builds
#2028 · EPNW-Eric · 2026-03-15
Issue with OGA granite-4.0-1b model response
#2020 · sanket131192 · 2026-03-12
Feature Request: Support for Qwen3.5
#2016 · Fazzioni · 2026-03-10
FP16 model supported for CPU
#2009 · bohuazou · 2026-03-05