Commits · 6bd0a983cd2cf74f27df2e5a5c80f1794a2ed7ef · OpenDAS / ollama

03 Apr, 2025 2 commits

model: support for mistral-small in the ollama runner · 6bd0a983

Bruce MacDonald authored Mar 14, 2025

Mistral is a popular research lab making open source models. This updates
the forward pass of llama architecture models to support both llama models
and mistral models by accounting for additional metadata present in mistral
models, and finding the correct dimensions for the output projection.

6bd0a983

fs: move ml.Config to fs package · 3b96a936
Michael Yang authored Mar 18, 2025

3b96a936

27 Mar, 2025 1 commit

ml: Remove Output from Context interface · 01aa7887

Jesse Gross authored Mar 27, 2025

Model implementations should use Input for all of their tensors
supplied to the model. This includes tensors that relate to the
outputs, which is confusing since there is also an Output funciton.

Since Output is only used internally in GGML and not used by any
model implementations, we can remove it from the interface to
reduce confusion.

01aa7887

21 Mar, 2025 1 commit
- ml/backend/ggml: load tensors in 32KiB chunks · 74bd0965
  Michael Yang authored Mar 19, 2025
  
  74bd0965
18 Mar, 2025 1 commit

ggml: return error on failure to read tensor data (#9872) · df94175a

Bruce MacDonald authored Mar 18, 2025

When converting a ggml model if there is a failure to read tensor data a nil error value was being returned. It should be assigned to the actual error from reading.

df94175a

17 Mar, 2025 2 commits
- ml/backend/ggml: allocate memory with malloc when loading model (#9822) · 364629b8
  Jeffrey Morgan authored Mar 17, 2025
  
  364629b8
- conditionally enable parallel pipelines · 4561fff3
  Michael Yang authored Mar 14, 2025
  
  4561fff3
11 Mar, 2025 8 commits
- use 2d pooling · 63a39406
  Michael Yang authored Mar 11, 2025
  
  63a39406
- fallback to cpu · c5cbe4fc
  Michael Yang authored Mar 10, 2025
  
  c5cbe4fc
- ollama debug tensor · 9e4642e9
  Michael Yang authored Mar 09, 2025
  
  9e4642e9
- duplicate token_embd to output · 6b0486c2
  Michael Yang authored Mar 09, 2025
  
  6b0486c2
- use fast attention · 8934324b
  Michael Yang authored Mar 07, 2025
  
  8934324b
- set non-causal attention · 0df18004
  Michael Yang authored Mar 07, 2025
  
  0df18004
- add gemma vision encoder · 4b037a97
  Michael Yang authored Mar 06, 2025
  
  4b037a97
- gemma2 impl · 5f74d1fd
  Patrick Devine authored Feb 07, 2025
  
  5f74d1fd
08 Mar, 2025 2 commits

ml: Add support for quantized KV cache · 4100ed7b

Jesse Gross authored Feb 21, 2025

Similar to the llama engine, quantizing the KV cache requires
flash attention to be enabled through the Ollama server.

4100ed7b

ggml-backend: Ensure allocation meet backend requirements · 25f9b152

Jesse Gross authored Mar 07, 2025

Backends can impose additional alignment requirements on buffer sizes.
We should ensure that we meet these or allocations can fail.

25f9b152

07 Mar, 2025 12 commits
- additional review comments · 98272fbd
  Jesse Gross authored Mar 07, 2025
  
  98272fbd
- ml/backend/ggml: use backend buffer type · b27e8f3f
  Michael Yang authored Mar 05, 2025
```
this ensures the tensor is created on the right buffer type for backends
such as cpu
```
  b27e8f3f
- comments · 45df786f
  Michael Yang authored Mar 04, 2025
  
  45df786f
- ml/backend/ggml: clean up · daaf42e4
  Michael Yang authored Feb 28, 2025
  
  daaf42e4
- ml/backend/ggml: offload vision to cpu · 2dc60d46
  Michael Yang authored Feb 27, 2025
```
temporary until tensor loading can accurately account for vision models
```
  2dc60d46
- ml/backend/ggml: handle tensor split · b5312f30
  Michael Yang authored Feb 26, 2025
  
  b5312f30
- ml/backend/ggml: handle user specified cpu offloading · 26c2e0bd
  Michael Yang authored Feb 26, 2025
  
  26c2e0bd
- ml/backend/ggml: set cpu n_threads · bf920883
  Michael Yang authored Feb 26, 2025
  
  bf920883
- ml/backend/ggml: create tensor on specific backend · 7bae7fa5
  Michael Yang authored Feb 25, 2025
```
some tensors should be created on specific backends to reduce number of
copies and improve performance
```
  7bae7fa5
- kvcache: create cache ctx per layer · 764e199d
  Michael Yang authored Feb 25, 2025
```
each cache layer creates and maintains its own context instead of using
a large context for all layers
```
  764e199d
- model: load non-repeated tensors into multiple backends · bfce55db
  Michael Yang authored Feb 24, 2025
```
some tensors are expected to be used in repeating layers but are not
themselves repeated. this change copies these tensors into the same
backends as their repeating counterparts to minimize copying tensors
between backends
```
  bfce55db
- ml/backend/ggml: update model loading for hybrid/multi backends · bab6f34d
  Michael Yang authored Feb 19, 2025
```
use a similar strategy as llama.cpp for deciding where tensors should be
allocated. this will be improved later to be aware of usable memory
before assigning the tensor
```
  bab6f34d
04 Mar, 2025 1 commit

ml/backend/ggml: consolidate system info logging · 05a01fde

Michael Yang authored Feb 28, 2025

- output backend system info when initializing the backend. this ensures
  this information is always present without needing to be called
  explicitly
- convert to structured logging
- enumerate devices rather than backends since devices are ordered
- track device indices grouped by device name

05a01fde

02 Mar, 2025 4 commits

ml: Enable support for flash attention · 21aa666a

Jesse Gross authored Feb 25, 2025

The GGML flash attention kernel has specific requirements for
padding and permutation. This adds support to the KV cache
for conforming to these requirements so that flash attention
can be enabled.

Flash attention can be used in the same situations as the llama
engine and is enabled by the user in the same way.

21aa666a

ml: Empty tensor constructor for tensors · ee141cc8

Jesse Gross authored Feb 28, 2025

In cases where we allocate a tensor and then fully overwrite it with
copied data, it is wasteful to first zero out the memory.

ee141cc8

ggml-backend: Store parent backend as part of tensor · 55e5776c

Jesse Gross authored Feb 27, 2025

It can be important for a tensor to know what backend it came from -
for example, to know if flash attention is enabled.

55e5776c

attention: Remove unnecessary contiguous operations · 854a9195

Jesse Gross authored Feb 22, 2025

Prior to performing attention, we need to permute query, key
and value. Currently we call Contiguous after each of these
permutations, which is correct but expensive. Avoiding the
3 calls to Contiguous increases performance by over 20%.

The permutations of query and key do not violate the continuity
rules for mulmat and the Contiguous call can be simply removed.

Value requires a different permutation and does require Contiguous.
However, we can use the copy into the cache as a way to perform this
without further overhead.

To support this and avoid unexpected tensor shapes that are seen by
models, we need tighter integration between attention, cache
and backend. Future optimization will also likely need this structure
 - for example, flash attention has special padding requirements in
the cache and other backends may have their own needs.

This further contains the operations that go into attention so that
these and other optimizations can be handled transparently. Models
that have special requirements for attention can still implement
their own version of it.

854a9195

27 Feb, 2025 1 commit

ml: update Context.Forward interface · 3e8b8a19

Michael Yang authored Feb 21, 2025

update Context.Forward to accept multiple tensors to match
Context.Compute signature

update Context.Forward to return Context such that it can be chained
with Context.Compute

3e8b8a19

21 Feb, 2025 2 commits

ml: Abstract attention out of model definitions · f53f4198

Jesse Gross authored Feb 14, 2025



There are two benefits to doing this:
 - Provide a library function that models can use, reducing code for
   each model implementation
 - Enables a single place to drop in optimized implementations of
   attention based on the backend or other factors. One is provided for
   GGML.

On CUDA this improves token generation rate by about 3%. It does not
have a significant effect on Metal.
Co-authored-by: Daniel Hiltgen <daniel@ollama.com>

f53f4198

ml/backend/ggml: fix rms norm · 2192a28e
Michael Yang authored Feb 20, 2025

2192a28e

20 Feb, 2025 2 commits

ggml-backend: Don't recreate the scheduler for each context · e5bcc51a

Jesse Gross authored Feb 18, 2025

We don't need to create and destroy the GGML scheduler for every
context. This introduces extra CPU overhead for every forward
pass and extra memory for contexts that don't actually get scheduled
(for example, KV caches). We can instead just have one scheduler
for the backend and reset it each time we call Compute.

This improves token generation performance by 1-2% and removes
scheduler create/destroy from profile traces.

e5bcc51a

ollamarunner: Pass runner performance parameters to backends · bd6a7d5e

Jesse Gross authored Feb 20, 2025

Currently the following parameters are in the runner but not used:
 - numGPULayers
 - mainGPU
 - threads
 - tensorSplit

This passes them through to the backend, which is where they would
actually get used. However, the GGML backend does not yet do anything
with them.

bd6a7d5e

14 Feb, 2025 1 commit
- Wire up system info log for new engine (#9123) · df2680b4
  Daniel Hiltgen authored Feb 14, 2025
  
  df2680b4