Benchmark, profile, and optimize your local LLM inference on consumer GPUs. Get token/sec gains without the cloud costs.
See actual token/sec rates on RTX 4090, 5080, 3090, and consumer cards. No synthetic numbers.
Batch size, quantization level, context window tuning. Automated recommendations based on your hardware.
Stop trying random flags. Profiler shows exactly where your bottleneck is: compute, memory, or bandwidth.
Run Qwen, Llama, Mixtral side-by-side on YOUR GPU. See which model actually fits your use case.
Collected from community runs on actual hardware:
| Model | GPU | Batch Size | Quant | Tokens/sec | VRAM Used |
|---|---|---|---|---|---|
| Llama 70B | RTX 5080 | 4 | Q4_K_M | 92 tok/sec | 38 GB |
| Qwen 27B | RTX 4090 | 8 | Q5_K_M | 156 tok/sec | 24 GB |
| Mixtral 8x7B | RTX 3090 (2x) | 2 | BF16 | 48 tok/sec | 45 GB |
| Llama 8B | RTX 4070 | 16 | Q3_K_S | 204 tok/sec | 12 GB |
Three steps to optimal inference:
Then open http://localhost:8080 to see real-time profiling and tuning recommendations.
Built for developers squeezing more tokens per dollar from their GPUs.
View on GitHub Read the DocsRegister interest
This is not a purchase and there is no card field. It puts your address, this product, and whatever you write below in front of a person, and you get a written answer about what finishing it, or handing it over for you to run yourself, would actually take.