Batch throughput

August 23, 2026 ยท View on GitHub

Recommended use
Start withA model-and-card pair with a current batch receipt
PrioritizeRequest mix, output length, admission limits, and aggregate completion rate
ValidateThe real HTTP surface at the intended concurrency, not only a kernel benchmark
Read nextServing, Performance, Testing

Do not assume the fastest single-stream path is the best batch path. Memra selects and qualifies decode modes per request shape; use the board and the server gate for the workload you will run.