llama.cpp

History

Michael Podvitskiy fb4a0ec083 llama : propagate the results of `graph_compute` (#9525 ) * llama: propagating the results of `graph_compute` to the user interface * llama: reverting kv_cache in case of failed compute * llama: `llama_kv_cache_state` was removed, only the result of `llama_graph_compute` is returned * llama: restore a kv_cache in case of failed computation * llama: correct reverting of the entire batch. also updates `llama_kv_cache_find_slot`, will correctly count the number of `used` cells for recurrent models * llama: updated comments * llama : add comments about KV cache state after error --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>		2024-11-13 20:00:35 +02:00
..
CMakeLists.txt	llama : move vocab, grammar and sampling into separate files (#8508 )	2024-07-23 13:10:17 +03:00
llama-grammar.cpp	llama : refactor sampling v2 (#9294 )	2024-09-07 15:16:19 +03:00
llama-grammar.h	llama : refactor sampling v2 (#9294 )	2024-09-07 15:16:19 +03:00
llama-impl.h	log : add CONT level for continuing previous log entry (#9610 )	2024-09-24 10:15:35 +03:00
llama-sampling.cpp	DRY: Fixes clone functionality (#10192 )	2024-11-07 16:20:25 +01:00
llama-sampling.h	llama : add DRY sampler (#9702 )	2024-10-25 19:07:34 +03:00
llama-vocab.cpp	llama : add DRY sampler (#9702 )	2024-10-25 19:07:34 +03:00
llama-vocab.h	llama : add DRY sampler (#9702 )	2024-10-25 19:07:34 +03:00
llama.cpp	llama : propagate the results of `graph_compute` (#9525 )	2024-11-13 20:00:35 +02:00
unicode-data.cpp	server : better security control for public deployments (#9776 )	2024-10-08 13:27:04 +02:00
unicode-data.h	llama : reduce compile time and binary size (#9712 )	2024-10-02 15:49:55 +02:00
unicode.cpp	llama : reduce compile time and binary size (#9712 )	2024-10-02 15:49:55 +02:00
unicode.h	llama : move vocab, grammar and sampling into separate files (#8508 )	2024-07-23 13:10:17 +03:00