batched : add bench tool (#3545)

* batched : add bench tool

* batched : minor fix table

* batched-bench : add readme + n_kv_max is now configurable

* batched-bench : init warm-up batch

* batched-bench : pass custom set of PP, TG and PL

* batched-bench : add mmq CLI arg
This commit is contained in:
Georgi Gerganov 2023-10-11 21:25:33 +03:00 committed by GitHub
parent 24ba3d829e
commit 8c70a5ff25
No known key found for this signature in database
GPG key ID: 4AEE18F83AFDEB23
7 changed files with 321 additions and 3 deletions

1
.gitignore vendored
View file

@ -55,6 +55,7 @@ models-mnt
/server
/simple
/batched
/batched-bench
/export-lora
/finetune
/speculative