skip to content
open source project

openinfer-project/openinfer

Credential-building open issues here — a merged pull request is verifiable proof of work for your résumé.

1
winnable issues
0%
merge velocity signal
602
31 contributors

What it is

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

Why get involved

Real code, on a real project, reviewed by a real maintainer — a merged PR here is credential-building proof of work you can point to.

Open issues

Check against my GitHub

Computed from your public GitHub only — nothing is stored.

Connect GitHub to check →

qwen3: PerToken decode-graph memory grows linearly with batch bucket and is not budgeted

openinfer-project/openinfer · issue #780 · opened Jul 2026 · no PRs referenced when indexed

help wantedqwen3hw:1-gpu

Under the PerToken batch-invariance policy, each bucket-N decode CUDA graph bakes ~36 layers × 6 projections × N cublasGemmEx nodes, because the PerToken GEMM…

How to jump in

Click Claim →on any issue above to open its claim page, where you'll get a terminalhire claim <token> command and a one-tap deep link into your terminal.

terminalhire contribute

Why we picked it

Winnable volume0.13
Merge velocity0.00
Freshness1.00
Skill coverage0.00
Popularity0.41