qwen3: PerToken decode-graph memory grows linearly with batch bucket and is not budgeted
openinfer-project/openinfer · issue #780 · opened Jul 2026 · no PRs referenced when indexed
help wantedqwen3hw:1-gpu
Under the PerToken batch-invariance policy, each bucket-N decode CUDA graph bakes ~36 layers × 6 projections × N cublasGemmEx nodes, because the PerToken GEMM…
Click Claim →on any issue above to open its claim page, where you'll get a terminalhire claim <token> command and a one-tap deep link into your terminal.