llama : add thread safety test (#14035)

* llama : add thread safety test * llamafile : remove global state * llama : better LLAMA_SPLIT_MODE_NONE logic when main_gpu < 0 GPU devices are not used --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2025-06-16 08:11:43 -07:00 · 2025-06-16 08:11:43 -07:00 · 6adc3c3ebc
commit 6adc3c3ebc
parent 0dbcabde8c
9 changed files with 192 additions and 18 deletions
--- a/tests/CMakeLists.txt
+++ b/tests/CMakeLists.txt
@ -185,6 +185,8 @@ llama_build_and_test(test-json-partial.cpp)
 llama_build_and_test(test-log.cpp)
 llama_build_and_test(test-regex-partial.cpp)

+llama_build_and_test(test-thread-safety.cpp ARGS -hf ggml-org/models -hff tinyllamas/stories15M-q4_0.gguf -ngl 99 -p "The meaning of life is" -n 128 -c 256 -ub 32 -np 4)
+
 # this fails on windows (github hosted runner) due to curl DLL not found (exit code 0xc0000135)
 if (NOT WIN32)
    llama_build_and_test(test-arg-parser.cpp)