llama.cpp

History

0cc4m ee1628bdfe Basic Vulkan Multi-GPU implementation (#5321 ) * Initial Vulkan multi-gpu implementation Move most global variables into backend context * Add names to backend device functions * Add further missing cleanup code * Reduce code duplication in tensor split layer assignment * generalize LLAMA_SPLIT_LAYER for all backends, do not expose device count and memory in llama.h * Only do device info print in the beginning and initialize one backend for cpu assist Add missing cleanup code * Rework backend memory management to make sure devices and buffers get properly allocated and freed * Rename cpu assist free function --------- Co-authored-by: slaren <slarengh@gmail.com>		2024-02-07 07:54:50 +01:00
..
base64.hpp	llava : expose as a shared library for downstream projects (#3613 )	2023-11-07 00:36:23 +03:00
build-info.cpp.in	build : link against build info instead of compiling against it (#3879 )	2023-11-02 08:50:16 +02:00
CMakeLists.txt	cmake : fix ld warning duplicate libraries libllama.a (#4671 )	2023-12-29 16:39:15 +02:00
common.cpp	Basic Vulkan Multi-GPU implementation (#5321 )	2024-02-07 07:54:50 +01:00
common.h	YaRN : store rope scaling type as int32_t in memory (#5285 )	2024-02-03 13:22:06 +02:00
console.cpp	check C++ code with -Wmissing-declarations (#3184 )	2023-09-15 15:38:27 -04:00
console.h	gguf : new file format with flexible meta data (beta) (#2398 )	2023-08-21 23:07:43 +03:00
grammar-parser.cpp	grammar-parser : fix typo (#4318 )	2023-12-04 09:57:35 +02:00
grammar-parser.h	gguf : new file format with flexible meta data (beta) (#2398 )	2023-08-21 23:07:43 +03:00
log.h	english : use `typos` to fix comments and logs (#4354 )	2023-12-12 11:53:36 +02:00
sampling.cpp	Remove unused data and add fixes (#5154 )	2024-01-27 15:25:55 +01:00
sampling.h	llama : dynamic temperature sampling (#4972 )	2024-01-25 22:06:22 +02:00
stb_image.h	examples: support LLaVA v1.5 (multimodal model) (#3436 )	2023-10-12 18:23:18 +03:00
train.cpp	llama : remove LLAMA_MAX_DEVICES and LLAMA_SUPPORTS_GPU_OFFLOAD (#5240 )	2024-01-31 17:30:17 +02:00
train.h	sync : ggml (backend v2) (#3912 )	2023-11-13 14:16:23 +02:00