Commits · 9c5bf342bc34f94de9aa4a171d726e6b341a91e6 · OpenDAS / ollama

17 Sep, 2025 4 commits
- fix: multi-cuda version skew (#12318) · 9c5bf342
  Daniel Hiltgen authored Sep 17, 2025
```
Ensure that in a version skewed multi-cuda setup we use the lowest version for all GPUs
```
  9c5bf342
- fix(llama): other llama flavours (#12308) · 564b558c
  Michael Yang authored Sep 17, 2025
```
* fix(llama): rope scale

* spm llama

* skip moe models

* cleanup
```
  564b558c
- prefer ollama engine for qwen3 (#12310) · a417ac97
  Michael Yang authored Sep 17, 2025
  
  a417ac97
- refactor: use the built-in max/min to simplify the code (#12280) · 05d53457
  russcoss authored Sep 16, 2025
```
Signed-off-by: russcoss <russcoss@outlook.com>
```
  05d53457
16 Sep, 2025 5 commits
- logutil: fix source field (#12279) · b225508c
  Michael Yang authored Sep 16, 2025
  
  b225508c
- Merge pull request #12248 from ollama/drifkin/qwen3-coder-parsing · fa1c987a
  Devon Rifkin authored Sep 16, 2025
```
add qwen3-coder tool support
```
  fa1c987a
- use split activations when possible (#12293) · ad95d5b3
  Michael Yang authored Sep 16, 2025
```
* use ggml_*_split activations when possible

* forward qkv
```
  ad95d5b3
- embed: cleanup (#12299) · c253433d
  Michael Yang authored Sep 16, 2025
```
* cleanup

* use pooling.TypeNone

* pooling test
```
  c253433d
- fix: fix CUDA detection for older GPUs (#12300) · a1cff89b
  Beshoy Girgis authored Sep 16, 2025
```
Prioritize GPU compute capability over driver version to ensure
Pascal GPUs (CC 6.1) use compatible CUDA v12 libraries instead of v13.
```
  a1cff89b
15 Sep, 2025 5 commits

doc: show how to clear the cgo cache (#12298) · 93c64ea1
Daniel Hiltgen authored Sep 15, 2025

93c64ea1

model: implement bert in ollama engine (#9080) · 3f6642f6

Michael Yang authored Sep 15, 2025

* fix truncate

* s/SentencePieceModel/SentencePiece/

* bert

* wordpiece

* refactor pooling

* more tokenizers

* normalize embeddings

3f6642f6

batch: use tensors for outputs (#12185) · 6f711714
Michael Yang authored Sep 15, 2025
```
this cleans up the model interface slightly without too much impact in
other areas
```
6f711714
address comments · 472feec2
Devon Rifkin authored Sep 15, 2025

472feec2

add qwen3-coder tool support · 47991940

Devon Rifkin authored Sep 11, 2025

The format qwen3-coder uses is relatively unique, both in rendering and
in parsing. To implement parsing, I wrote a custom parser in similar
style to harmony. For the rendering, I found that the logic would be
much more difficult to follow in a template, so I introduced the concept
of a built-in renderer that uses go code, rather than a template to
generate prompts.

I set us up for future built-in parsers and renderers by making it so
they can be specified in a Modelfile like so:

```
RENDERER "qwen3-coder"
PARSER "qwen3-coder"
```

These need to be provided explicitly because the architecture alone is
not enough to understand what format the model expects to receive, and
what format we expect it to output (e.g., qwen3-coder is `qwen3moe`,
which includes other qwen3-family models as well)

I haven't converted harmony to be one of these "built-ins" yet, since
some of it is in flux with the changes @ParthSareen has been making to
move harmony to the runner. It is likely that many other built-ins will
need to move to the runner as well, but I'm able to slightly defer that
decision since qwen3-coder doesn't have thinking (and therefore doesn't
need to be in the runner to make structured outputs work). I expect to
unify harmony with this approach very soon.

Whether a particular model supports tools or thinking was previously
inferred from templates, but without a template we now also use the
parser itself to declare what it supports. If we have future models that
re-use the same parsing format, but have different capabilities, we'll
want to parameterize them and give them different names to be specified
as a `PARSER`.

Misc changes:

- I worked on the renderer by diffing outputs from the reference
  implementation and ours. To make it easier to do this, I extended
  <https://github.com/ollama/ollama/pull/11875> to also support
  returning the prompt via the openai compat layer

47991940

12 Sep, 2025 5 commits

Revert "runner: move harmony to runner (#12052)" · 92b96d54
jmorganca authored Sep 12, 2025
```
This reverts commit 1a558f98.
```
92b96d54
Revert "runner: simplify parser entrypoints in runner (#12233)" · 9d56e63d
jmorganca authored Sep 12, 2025
```
This reverts commit 8d6fffae.
```
9d56e63d

Fix image cannot be seen with slice image on llama engine · 05309218

tc-mb authored Sep 13, 2025

Ollama's recent engine update, llama.cpp, caused all models requiring a slice schema to not display images. As a result, the value of numTokens isn't always the length of the sliced image embed, but rather the end length of the schema. This causes the image embed to not be correctly included during all slice processing.

05309218

tests: tighten up a few flaky tests (#12271) · 44a67928

Daniel Hiltgen authored Sep 12, 2025

Sometimes the context test results are pure emoji's
Thanksgiving has too much variability, so swap for a more straight forward prompt.

44a67928

cuda: remove compression for better compatibility (#12259) · e4ce6831
Daniel Hiltgen authored Sep 12, 2025
```
This retains compatibility with driver 531 and up at the trade-off of space.
```
e4ce6831

11 Sep, 2025 6 commits

ollamarunner: Suppress stack trace during memory allocation · 26214125

Jesse Gross authored Sep 11, 2025

Allocation failures can be a normal part of new memory estimates, so
we shouldn't print a stack trace in this case.

26214125

CI: fix windows cuda build (#12246) · 61fb912c

Daniel Hiltgen authored Sep 11, 2025

* ci: adjust cuda component list

v13 has a different breakdown of the components required to build ollama

* review comments

61fb912c

llm: Don't try to load split vision models in the Ollama engine · aba15753

Jesse Gross authored Sep 10, 2025

If a model with a split vision projector is loaded in the Ollama
engine, the projector will be ignored and the model will hallucinate
a response. Instead, fallback and try to load the model in the llama
engine.

aba15753

llm: Enable new memory estimates by default · eb10390d

Jesse Gross authored Sep 11, 2025

New memory estimates (see #11090 for more information) are now
enabled automatically for all models running on the Ollama engine,
improving both stability and performance through more accurate sizing
and allocation. Models running on the llama engine will continue to
use the original style of memory estimation.

eb10390d

feat: add dimensions field to embed requests (#12242) · feb18cd7
Michael Yang authored Sep 11, 2025
```
* feat: add field to truncate embeddings

* add openai embeddings for dimensions
```
feb18cd7
cmd: use slices.Contains to simplify code (#12249) · 8a7e2055
fengyuchuanshen authored Sep 12, 2025

8a7e2055

10 Sep, 2025 5 commits

ggml: Disable flash attention for gemma2 · 29ddfc2c

Jesse Gross authored Sep 09, 2025

Our new engine implementation of gemma2 doesn't support flash
attention, which means that it also doesn't support KV cache
quantization. Currently, it is possible to turn these two on,
which will result in a crash.

29ddfc2c

llm: Remove unneeded warning with flash attention enabled · 71cb86af

Jesse Gross authored Sep 09, 2025

If flash attention is enabled without KV cache quanitization, we will
currently always get this warning:
level=WARN source=server.go:226 msg="kv cache type not supported by model" type=""

71cb86af

docs: add ollama-co2 to community integrations (#12230) · 51989563
CarbonatedWater.org authored Sep 10, 2025

51989563

Add v12 + v13 cuda support (#12000) · 17a023f3

Daniel Hiltgen authored Sep 10, 2025

* Add support for upcoming NVIDIA Jetsons

The latest Jetsons with JetPack 7 are moving to an SBSA compatible model and
will not require building a JetPack specific variant.

* cuda: bring back dual versions

This adds back dual CUDA versions for our releases,
with v11 and v13 to cover a broad set of GPUs and
driver versions.

* win: break up native builds in build_windows.ps1

* v11 build working on windows and linux

* switch to cuda v12.8 not JIT

* Set CUDA compression to size

* enhance manual install linux docs

17a023f3

runner: simplify parser entrypoints in runner (#12233) · 8d6fffae
Parth Sareen authored Sep 10, 2025

8d6fffae

09 Sep, 2025 4 commits

tests: add tool calling integration test (#12232) · 20b53eaa
Parth Sareen authored Sep 09, 2025

20b53eaa

tests: reduce stress on CPU to 2 models (#12161) · 67451828

Daniel Hiltgen authored Sep 09, 2025

* tests: reduce stress on CPU to 2 models

This should avoid flakes due to systems getting overloaded with 3 (or more) models running concurrently

* tests: allow slow systems to pass on timeout

If a slow system is still streaming a response, and the response
will pass validation, don't fail just because the system is slow.

* test: unload embedding models more quickly

67451828

readme: add Clueless to community integrations (#12188) · f810ec74
Kashyap Tanuku authored Sep 09, 2025

f810ec74

llm: Clamp batch size to context size · e119783e

Jesse Gross authored Sep 08, 2025

The context must always be able to store the current batch, so
if the user requests a small context then we should also shrink
the batch to match. This also fixes the TestLongInputContext
test on the new engine. (The old engine already has this behavior.)

e119783e

08 Sep, 2025 4 commits

runner: move harmony to runner (#12052) · 1a558f98
Parth Sareen authored Sep 08, 2025

1a558f98

Hybrid and recurrent memory estimates (#12186) · 7b91c9ce

Gabe Goodhart authored Sep 08, 2025

This PR updates the memory size estimate logic to better handle recurrent and hybrid-recurrent models which are currently being badly overestimated because the default logic assumes full attention for all layers.

The logic for the sizing of the recurrent layers comes from the llama.cpp implementation

ggml_tensor * r = ggml_new_tensor_1d(ctx, type_r, hparams.n_embd_r()*mem_size);
ggml_tensor * s = ggml_new_tensor_1d(ctx, type_s, hparams.n_embd_s()*mem_size);
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

7b91c9ce

docs: show how to debug nvidia init failures (#12216) · 950d33aa
Daniel Hiltgen authored Sep 08, 2025
```
This debug setting can help troubleshoot obscure initialization failures.
```
950d33aa
fix: nil pointer dereference if cache is nil (#12215) · 9714e38d
Michael Yang authored Sep 08, 2025

9714e38d

05 Sep, 2025 1 commit

parser: don't check the file type of safetensors to prevent false negatives. (#12176) · 4378ae4f

frob authored Sep 06, 2025



* Don't check the file type of safetensor to prevent false negatives.

---------
Co-authored-by: Patrick Devine <patrick@infrahq.com>

4378ae4f

04 Sep, 2025 1 commit
- embedding gemma model (#12181) · 5994e8e8
  Michael Yang authored Sep 04, 2025
```
* ollama: add embeddings
```
  5994e8e8