llama.cpp

History

Jesse Gross f057808ffa ggml: Don't assert fail when tensor data changes (#13222 ) The following scenario will cause an assertion failure in the graph allocator: - Build and allocate a graph containing a tensor with a non-NULL data pointer - Build and allocate a new graph where that data is NULL Result: ggml-alloc.c:819: GGML_ASSERT(talloc->buffer_id >= 0) failed This happens during revalidation because we think that memory should have been previously allocated based on the current graph but in reality the previous graph was different. In this situation, we should do a full reallocation pass.		2025-05-01 22:46:10 +02:00
..
ggml-blas	ggml : add support for dynamic loading of backends (#10469 )	2024-11-25 15:13:39 +01:00
ggml-cann	CANN: Add support for async operator submission (#12864 )	2025-04-17 20:34:16 +08:00
ggml-cpu	ggml : fix ppc64le build (#13176 )	2025-04-30 13:17:08 +02:00
ggml-cuda	build : fix build info on windows (#13239 )	2025-05-01 21:48:08 +02:00
ggml-hip	CUDA/HIP: Share the same unified memory allocation logic. (#12934 )	2025-04-15 11:20:38 +02:00
ggml-kompute	llama : add Qwen2VL support + multimodal RoPE (#10361 )	2024-12-14 14:43:46 +02:00
ggml-metal	metal : fix floating-point range of attention scores in FA kernels (#13090 )	2025-04-24 10:38:30 +03:00
ggml-musa	cuda : enable CUDA Graph on CUDA Toolkit < 12.x (#12394 )	2025-03-17 20:25:13 +02:00
ggml-opencl	opencl: fix incorrect local_size index in profiling log (#12868 )	2025-04-16 14:25:57 -07:00
ggml-rpc	fix(rpc): Improve input validation and error handling (#13069 )	2025-04-28 21:00:20 +03:00
ggml-sycl	SYCL: Add all missing unary kernels (#13074 )	2025-04-28 11:33:25 +02:00
ggml-vulkan	vulkan: Add bfloat16 support (#12554 )	2025-05-01 20:49:39 +02:00
CMakeLists.txt	ggml : add SSE 4.2 and x64 base variant for CPUs without AVX (#12871 )	2025-04-21 18:13:51 +02:00
ggml-alloc.c	ggml: Don't assert fail when tensor data changes (#13222 )	2025-05-01 22:46:10 +02:00
ggml-backend-impl.h	ggml : upgrade init_tensor API to return a ggml_status (#11854 )	2025-02-28 14:41:47 +01:00
ggml-backend-reg.cpp	ggml-backend : fix backend search path (#12330 )	2025-03-11 14:25:17 +01:00
ggml-backend.cpp	ggml : portability fixes for VS 2017 (#12150 )	2025-03-04 18:53:26 +02:00
ggml-common.h	musa: fix all warnings, re-enable `-DLLAMA_FATAL_WARNINGS=ON` in ci and update doc (#12611 )	2025-03-30 10:59:38 +02:00
ggml-impl.h	ggml: don't include arm_neon.h when using CUDA 12 with ARM Neon (ggml/1187)	2025-04-11 00:17:47 +03:00
ggml-opt.cpp	ggml-opt: fix data corruption (ggml/1022)	2024-11-21 09:22:02 +02:00
ggml-quants.c	ggml : portability fixes for VS 2017 (#12150 )	2025-03-04 18:53:26 +02:00
ggml-quants.h	ggml : build backends as libraries (#10256 )	2024-11-14 18:04:35 +01:00
ggml-threading.cpp	ggml : build backends as libraries (#10256 )	2024-11-14 18:04:35 +01:00
ggml-threading.h	remove CMAKE_WINDOWS_EXPORT_ALL_SYMBOLS (#10797 )	2024-12-12 19:02:49 +01:00
ggml.c	ggml: move fp16/bf16 conversion optimizations to CPU backend + export conversion APIs (#13107 )	2025-04-26 16:05:31 +02:00
gguf.cpp	Fix clang warning in gguf_check_reserved_keys (#12686 )	2025-04-01 13:12:53 +02:00