llama.cpp

History

Georgi Gerganov a19b5cef16 llama : fix FA when KV cache is not used (i.e. embeddings) (#12825 ) * ggml : FA supports F32 V * graph : cast KV to F16 when the KV cache is not used ggml-ci * server : add test that exercises embeddings with FA enabled ggml-ci		2025-04-08 19:54:51 +03:00
..
test_basic.py	server : add flag to disable the web-ui (#10762 ) (#10751 )	2024-12-10 18:22:34 +01:00
test_chat_completion.py	`server`: fix deadly typo in response_format.json_schema.schema handling (#12168 )	2025-03-04 08:24:07 +02:00
test_completion.py	server : Fixed wrong function name in llamacpp server unit test (#11473 )	2025-01-29 00:03:42 +01:00
test_ctx_shift.py	server : replace behave with pytest (#10416 )	2024-11-26 16:20:18 +01:00
test_embedding.py	llama : fix FA when KV cache is not used (i.e. embeddings) (#12825 )	2025-04-08 19:54:51 +03:00
test_infill.py	server : fix extra BOS in infill endpoint (#11106 )	2025-01-06 15:36:08 +02:00
test_lora.py	server : allow using LoRA adapters per-request (#10994 )	2025-01-02 15:05:18 +01:00
test_rerank.py	server : add TEI API format for /rerank endpoint (#11942 )	2025-02-18 14:21:41 +01:00
test_security.py	server : replace behave with pytest (#10416 )	2024-11-26 16:20:18 +01:00
test_slot_save.py	server : replace behave with pytest (#10416 )	2024-11-26 16:20:18 +01:00
test_speculative.py	server : allow using LoRA adapters per-request (#10994 )	2025-01-02 15:05:18 +01:00
test_tokenize.py	server : replace behave with pytest (#10416 )	2024-11-26 16:20:18 +01:00
test_tool_call.py	`tool-call`: ensure there's always a non-empty tool call id (#12292 )	2025-03-10 09:45:29 +00:00