Commits · af060eb2508e8bed25241163243bdd7471cb7fd6 · OpenDAS / ollama

22 Sep, 2025 2 commits
- docs: update cloud.md for cloud models · af060eb2
  jmorganca authored Sep 19, 2025
  
  af060eb2
- docs: move turbo.md to cloud.md · ae5c3300
  jmorganca authored Sep 19, 2025
  
  ae5c3300
20 Sep, 2025 2 commits

Merge pull request #12358 from ollama/drifkin/qwen3-coder-ampersands · 3677842f
Devon Rifkin authored Sep 20, 2025
```
parsers: fix `&`s in qwen3coder parameter values
```
3677842f

parsers: fix `&`s in qwen3coder parameter values · 242df70a

Devon Rifkin authored Sep 20, 2025

In <https://github.com/ollama/ollama/issues/12357> we that the model
will output tool calls such as

```
<function=shell>
<parameter=command>
pwd && ls -la
</parameter>
</function>
```

We parse this using the approach of transforming into valid xml and then
using an xml parser. While we do transform the function and parameter
names, we weren't escaping the parameter values (which in this example
are invalid since `pwd && ls -la` contains unescaped ampersands).

This has been fixed by first transforming the tags in the same way, and
then walking the transformed string and escaping the text in between the
tags. This also fixes a case where `<` in the middle of a parameter
value would cause an xml parse failure.

Fixes: #12357

242df70a

19 Sep, 2025 1 commit
- gemma: fix rope scaling for qat models (#12348) · dba39b2e
  Patrick Devine authored Sep 19, 2025
```
* gemma: fix rope scaling for qat models

* gofumpt yourself
```
  dba39b2e
18 Sep, 2025 7 commits

fix: model load for unsupported embedding models (#12311) · 9f3a37fd

Michael Yang authored Sep 18, 2025

with #12181, there's now support for embeddings in ollama engine.
this is done by mutating the architecture and adding _embed when it
detects an embedding model. however this introduced a bug where if
an embedding model was run based on an existing ollama engine model
without an embedding implementation, e.g. llama4, it will pass the
initial arch support check but fail when actually loaded.

there's currently two entrypoints to creating a model. previously this
second entrypoint was necessary because calling model.New would also
load the model. since #11818, this is no longer th case so merge them
to reduce complexity

9f3a37fd

feat: qwen3 embed (#12301) · 7460259e
Michael Yang authored Sep 18, 2025
```
* cleanup

* use pooling.TypeNone

* pooling test

* qwen3 embed
```
7460259e
server: add unauthorized error to remote chat handler (#12338) · 22ccdd74
Jeffrey Morgan authored Sep 18, 2025

22ccdd74

build: avoid unbounded parallel builds (#12319) · 0c3d0e75

Daniel Hiltgen authored Sep 18, 2025

With the addition of cuda v13, on a clean setup, the level of parallelism
was causing docker desktop to become overwhelmed and compilers
were crashing. This limits to 8 parallel per build stage, with the ability
to override if you have many more cores available.

0c3d0e75

auth: check the permissions on the private key to see if it's readable (#12336) · eb0a5d44
Patrick Devine authored Sep 18, 2025

eb0a5d44
fix(integration): check truncated length (#12337) · ceac416e
Michael Yang authored Sep 18, 2025

ceac416e

convert: convert bf16 vision weights to fp16 (#12324) · 2717dce6

Patrick Devine authored Sep 17, 2025

This change moves back to converting bf16 vision weights to fp16,
specifically if they start with the name "v." (such as v.blk.0.attn_k.weight).

This fixes a bug where converted images are failing because they are trying
to call `im2col` which doesn't have a bf16 kernel in ggml.

2717dce6

17 Sep, 2025 6 commits
- server: skip parsing initial <think> if provided in the prompt for /api/generate (#12289) · 9b8187b4
  frob authored Sep 18, 2025
  
  9b8187b4
- engine: add remote proxy (#12307) · 8b894933
  Patrick Devine authored Sep 17, 2025
  
  8b894933
- fix: multi-cuda version skew (#12318) · 9c5bf342
  Daniel Hiltgen authored Sep 17, 2025
```
Ensure that in a version skewed multi-cuda setup we use the lowest version for all GPUs
```
  9c5bf342
- fix(llama): other llama flavours (#12308) · 564b558c
  Michael Yang authored Sep 17, 2025
```
* fix(llama): rope scale

* spm llama

* skip moe models

* cleanup
```
  564b558c
- prefer ollama engine for qwen3 (#12310) · a417ac97
  Michael Yang authored Sep 17, 2025
  
  a417ac97
- refactor: use the built-in max/min to simplify the code (#12280) · 05d53457
  russcoss authored Sep 16, 2025
```
Signed-off-by: russcoss <russcoss@outlook.com>
```
  05d53457
16 Sep, 2025 5 commits
- logutil: fix source field (#12279) · b225508c
  Michael Yang authored Sep 16, 2025
  
  b225508c
- Merge pull request #12248 from ollama/drifkin/qwen3-coder-parsing · fa1c987a
  Devon Rifkin authored Sep 16, 2025
```
add qwen3-coder tool support
```
  fa1c987a
- use split activations when possible (#12293) · ad95d5b3
  Michael Yang authored Sep 16, 2025
```
* use ggml_*_split activations when possible

* forward qkv
```
  ad95d5b3
- embed: cleanup (#12299) · c253433d
  Michael Yang authored Sep 16, 2025
```
* cleanup

* use pooling.TypeNone

* pooling test
```
  c253433d
- fix: fix CUDA detection for older GPUs (#12300) · a1cff89b
  Beshoy Girgis authored Sep 16, 2025
```
Prioritize GPU compute capability over driver version to ensure
Pascal GPUs (CC 6.1) use compatible CUDA v12 libraries instead of v13.
```
  a1cff89b
15 Sep, 2025 5 commits

doc: show how to clear the cgo cache (#12298) · 93c64ea1
Daniel Hiltgen authored Sep 15, 2025

93c64ea1

model: implement bert in ollama engine (#9080) · 3f6642f6

Michael Yang authored Sep 15, 2025

* fix truncate

* s/SentencePieceModel/SentencePiece/

* bert

* wordpiece

* refactor pooling

* more tokenizers

* normalize embeddings

3f6642f6

batch: use tensors for outputs (#12185) · 6f711714
Michael Yang authored Sep 15, 2025
```
this cleans up the model interface slightly without too much impact in
other areas
```
6f711714
address comments · 472feec2
Devon Rifkin authored Sep 15, 2025

472feec2

add qwen3-coder tool support · 47991940

Devon Rifkin authored Sep 11, 2025

The format qwen3-coder uses is relatively unique, both in rendering and
in parsing. To implement parsing, I wrote a custom parser in similar
style to harmony. For the rendering, I found that the logic would be
much more difficult to follow in a template, so I introduced the concept
of a built-in renderer that uses go code, rather than a template to
generate prompts.

I set us up for future built-in parsers and renderers by making it so
they can be specified in a Modelfile like so:

```
RENDERER "qwen3-coder"
PARSER "qwen3-coder"
```

These need to be provided explicitly because the architecture alone is
not enough to understand what format the model expects to receive, and
what format we expect it to output (e.g., qwen3-coder is `qwen3moe`,
which includes other qwen3-family models as well)

I haven't converted harmony to be one of these "built-ins" yet, since
some of it is in flux with the changes @ParthSareen has been making to
move harmony to the runner. It is likely that many other built-ins will
need to move to the runner as well, but I'm able to slightly defer that
decision since qwen3-coder doesn't have thinking (and therefore doesn't
need to be in the runner to make structured outputs work). I expect to
unify harmony with this approach very soon.

Whether a particular model supports tools or thinking was previously
inferred from templates, but without a template we now also use the
parser itself to declare what it supports. If we have future models that
re-use the same parsing format, but have different capabilities, we'll
want to parameterize them and give them different names to be specified
as a `PARSER`.

Misc changes:

- I worked on the renderer by diffing outputs from the reference
  implementation and ours. To make it easier to do this, I extended
  <https://github.com/ollama/ollama/pull/11875> to also support
  returning the prompt via the openai compat layer

47991940

12 Sep, 2025 5 commits

Revert "runner: move harmony to runner (#12052)" · 92b96d54
jmorganca authored Sep 12, 2025
```
This reverts commit 1a558f98.
```
92b96d54
Revert "runner: simplify parser entrypoints in runner (#12233)" · 9d56e63d
jmorganca authored Sep 12, 2025
```
This reverts commit 8d6fffae.
```
9d56e63d

Fix image cannot be seen with slice image on llama engine · 05309218

tc-mb authored Sep 13, 2025

Ollama's recent engine update, llama.cpp, caused all models requiring a slice schema to not display images. As a result, the value of numTokens isn't always the length of the sliced image embed, but rather the end length of the schema. This causes the image embed to not be correctly included during all slice processing.

05309218

tests: tighten up a few flaky tests (#12271) · 44a67928

Daniel Hiltgen authored Sep 12, 2025

Sometimes the context test results are pure emoji's
Thanksgiving has too much variability, so swap for a more straight forward prompt.

44a67928

cuda: remove compression for better compatibility (#12259) · e4ce6831
Daniel Hiltgen authored Sep 12, 2025
```
This retains compatibility with driver 531 and up at the trade-off of space.
```
e4ce6831

11 Sep, 2025 6 commits

ollamarunner: Suppress stack trace during memory allocation · 26214125

Jesse Gross authored Sep 11, 2025

Allocation failures can be a normal part of new memory estimates, so
we shouldn't print a stack trace in this case.

26214125

CI: fix windows cuda build (#12246) · 61fb912c

Daniel Hiltgen authored Sep 11, 2025

* ci: adjust cuda component list

v13 has a different breakdown of the components required to build ollama

* review comments

61fb912c

llm: Don't try to load split vision models in the Ollama engine · aba15753

Jesse Gross authored Sep 10, 2025

If a model with a split vision projector is loaded in the Ollama
engine, the projector will be ignored and the model will hallucinate
a response. Instead, fallback and try to load the model in the llama
engine.

aba15753

llm: Enable new memory estimates by default · eb10390d

Jesse Gross authored Sep 11, 2025

New memory estimates (see #11090 for more information) are now
enabled automatically for all models running on the Ollama engine,
improving both stability and performance through more accurate sizing
and allocation. Models running on the llama engine will continue to
use the original style of memory estimation.

eb10390d

feat: add dimensions field to embed requests (#12242) · feb18cd7
Michael Yang authored Sep 11, 2025
```
* feat: add field to truncate embeddings

* add openai embeddings for dimensions
```
feb18cd7
cmd: use slices.Contains to simplify code (#12249) · 8a7e2055
fengyuchuanshen authored Sep 12, 2025

8a7e2055

10 Sep, 2025 1 commit

ggml: Disable flash attention for gemma2 · 29ddfc2c

Jesse Gross authored Sep 09, 2025

Our new engine implementation of gemma2 doesn't support flash
attention, which means that it also doesn't support KV cache
quantization. Currently, it is possible to turn these two on,
which will result in a crash.

29ddfc2c