Commits · 472feec2ff5096eb23f72356f26d67b71f18d01e · OpenDAS / ollama

15 Sep, 2025 2 commits

address comments · 472feec2
Devon Rifkin authored Sep 15, 2025

472feec2

add qwen3-coder tool support · 47991940

Devon Rifkin authored Sep 11, 2025

The format qwen3-coder uses is relatively unique, both in rendering and
in parsing. To implement parsing, I wrote a custom parser in similar
style to harmony. For the rendering, I found that the logic would be
much more difficult to follow in a template, so I introduced the concept
of a built-in renderer that uses go code, rather than a template to
generate prompts.

I set us up for future built-in parsers and renderers by making it so
they can be specified in a Modelfile like so:

```
RENDERER "qwen3-coder"
PARSER "qwen3-coder"
```

These need to be provided explicitly because the architecture alone is
not enough to understand what format the model expects to receive, and
what format we expect it to output (e.g., qwen3-coder is `qwen3moe`,
which includes other qwen3-family models as well)

I haven't converted harmony to be one of these "built-ins" yet, since
some of it is in flux with the changes @ParthSareen has been making to
move harmony to the runner. It is likely that many other built-ins will
need to move to the runner as well, but I'm able to slightly defer that
decision since qwen3-coder doesn't have thinking (and therefore doesn't
need to be in the runner to make structured outputs work). I expect to
unify harmony with this approach very soon.

Whether a particular model supports tools or thinking was previously
inferred from templates, but without a template we now also use the
parser itself to declare what it supports. If we have future models that
re-use the same parsing format, but have different capabilities, we'll
want to parameterize them and give them different names to be specified
as a `PARSER`.

Misc changes:

- I worked on the renderer by diffing outputs from the reference
  implementation and ours. To make it easier to do this, I extended
  <https://github.com/ollama/ollama/pull/11875> to also support
  returning the prompt via the openai compat layer

47991940

12 Sep, 2025 5 commits

Revert "runner: move harmony to runner (#12052)" · 92b96d54
jmorganca authored Sep 12, 2025
```
This reverts commit 1a558f98.
```
92b96d54
Revert "runner: simplify parser entrypoints in runner (#12233)" · 9d56e63d
jmorganca authored Sep 12, 2025
```
This reverts commit 8d6fffae.
```
9d56e63d

Fix image cannot be seen with slice image on llama engine · 05309218

tc-mb authored Sep 13, 2025

Ollama's recent engine update, llama.cpp, caused all models requiring a slice schema to not display images. As a result, the value of numTokens isn't always the length of the sliced image embed, but rather the end length of the schema. This causes the image embed to not be correctly included during all slice processing.

05309218

tests: tighten up a few flaky tests (#12271) · 44a67928

Daniel Hiltgen authored Sep 12, 2025

Sometimes the context test results are pure emoji's
Thanksgiving has too much variability, so swap for a more straight forward prompt.

44a67928

cuda: remove compression for better compatibility (#12259) · e4ce6831
Daniel Hiltgen authored Sep 12, 2025
```
This retains compatibility with driver 531 and up at the trade-off of space.
```
e4ce6831

11 Sep, 2025 6 commits

ollamarunner: Suppress stack trace during memory allocation · 26214125

Jesse Gross authored Sep 11, 2025

Allocation failures can be a normal part of new memory estimates, so
we shouldn't print a stack trace in this case.

26214125

CI: fix windows cuda build (#12246) · 61fb912c

Daniel Hiltgen authored Sep 11, 2025

* ci: adjust cuda component list

v13 has a different breakdown of the components required to build ollama

* review comments

61fb912c

llm: Don't try to load split vision models in the Ollama engine · aba15753

Jesse Gross authored Sep 10, 2025

If a model with a split vision projector is loaded in the Ollama
engine, the projector will be ignored and the model will hallucinate
a response. Instead, fallback and try to load the model in the llama
engine.

aba15753

llm: Enable new memory estimates by default · eb10390d

Jesse Gross authored Sep 11, 2025

New memory estimates (see #11090 for more information) are now
enabled automatically for all models running on the Ollama engine,
improving both stability and performance through more accurate sizing
and allocation. Models running on the llama engine will continue to
use the original style of memory estimation.

eb10390d

feat: add dimensions field to embed requests (#12242) · feb18cd7
Michael Yang authored Sep 11, 2025
```
* feat: add field to truncate embeddings

* add openai embeddings for dimensions
```
feb18cd7
cmd: use slices.Contains to simplify code (#12249) · 8a7e2055
fengyuchuanshen authored Sep 12, 2025

8a7e2055

10 Sep, 2025 5 commits

ggml: Disable flash attention for gemma2 · 29ddfc2c

Jesse Gross authored Sep 09, 2025

Our new engine implementation of gemma2 doesn't support flash
attention, which means that it also doesn't support KV cache
quantization. Currently, it is possible to turn these two on,
which will result in a crash.

29ddfc2c

llm: Remove unneeded warning with flash attention enabled · 71cb86af

Jesse Gross authored Sep 09, 2025

If flash attention is enabled without KV cache quanitization, we will
currently always get this warning:
level=WARN source=server.go:226 msg="kv cache type not supported by model" type=""

71cb86af

docs: add ollama-co2 to community integrations (#12230) · 51989563
CarbonatedWater.org authored Sep 10, 2025

51989563

Add v12 + v13 cuda support (#12000) · 17a023f3

Daniel Hiltgen authored Sep 10, 2025

* Add support for upcoming NVIDIA Jetsons

The latest Jetsons with JetPack 7 are moving to an SBSA compatible model and
will not require building a JetPack specific variant.

* cuda: bring back dual versions

This adds back dual CUDA versions for our releases,
with v11 and v13 to cover a broad set of GPUs and
driver versions.

* win: break up native builds in build_windows.ps1

* v11 build working on windows and linux

* switch to cuda v12.8 not JIT

* Set CUDA compression to size

* enhance manual install linux docs

17a023f3

runner: simplify parser entrypoints in runner (#12233) · 8d6fffae
Parth Sareen authored Sep 10, 2025

8d6fffae

09 Sep, 2025 4 commits

tests: add tool calling integration test (#12232) · 20b53eaa
Parth Sareen authored Sep 09, 2025

20b53eaa

tests: reduce stress on CPU to 2 models (#12161) · 67451828

Daniel Hiltgen authored Sep 09, 2025

* tests: reduce stress on CPU to 2 models

This should avoid flakes due to systems getting overloaded with 3 (or more) models running concurrently

* tests: allow slow systems to pass on timeout

If a slow system is still streaming a response, and the response
will pass validation, don't fail just because the system is slow.

* test: unload embedding models more quickly

67451828

readme: add Clueless to community integrations (#12188) · f810ec74
Kashyap Tanuku authored Sep 09, 2025

f810ec74

llm: Clamp batch size to context size · e119783e

Jesse Gross authored Sep 08, 2025

The context must always be able to store the current batch, so
if the user requests a small context then we should also shrink
the batch to match. This also fixes the TestLongInputContext
test on the new engine. (The old engine already has this behavior.)

e119783e

08 Sep, 2025 4 commits

runner: move harmony to runner (#12052) · 1a558f98
Parth Sareen authored Sep 08, 2025

1a558f98

Hybrid and recurrent memory estimates (#12186) · 7b91c9ce

Gabe Goodhart authored Sep 08, 2025

This PR updates the memory size estimate logic to better handle recurrent and hybrid-recurrent models which are currently being badly overestimated because the default logic assumes full attention for all layers.

The logic for the sizing of the recurrent layers comes from the llama.cpp implementation

ggml_tensor * r = ggml_new_tensor_1d(ctx, type_r, hparams.n_embd_r()*mem_size);
ggml_tensor * s = ggml_new_tensor_1d(ctx, type_s, hparams.n_embd_s()*mem_size);
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

7b91c9ce

docs: show how to debug nvidia init failures (#12216) · 950d33aa
Daniel Hiltgen authored Sep 08, 2025
```
This debug setting can help troubleshoot obscure initialization failures.
```
950d33aa
fix: nil pointer dereference if cache is nil (#12215) · 9714e38d
Michael Yang authored Sep 08, 2025

9714e38d

05 Sep, 2025 1 commit

parser: don't check the file type of safetensors to prevent false negatives. (#12176) · 4378ae4f

frob authored Sep 06, 2025



* Don't check the file type of safetensor to prevent false negatives.

---------
Co-authored-by: Patrick Devine <patrick@infrahq.com>

4378ae4f

04 Sep, 2025 2 commits
- embedding gemma model (#12181) · 5994e8e8
  Michael Yang authored Sep 04, 2025
```
* ollama: add embeddings
```
  5994e8e8
- more logutil.Trace (#12177) · b3e61207
  Michael Yang authored Sep 03, 2025
  
  b3e61207
02 Sep, 2025 3 commits
- logutil: add Trace and TraceContext helpers (#12110) · fb92b617
  Michael Yang authored Sep 02, 2025
  
  fb92b617
- llm: Avoid underflow in free memory logging · 8149a3c8
  Jesse Gross authored Sep 02, 2025
```
If a GPU's free memory is less than the reserved amount, we might get
an underflow. Since it is an unsigned uint64, we print this as a large
number rather than the more correct 0. This only affects logging, the
actual layout code already handles this correctly.

Bug #12138
```
  8149a3c8
- harden uncaught exception registration (#12120) · 0cc90a81
  Daniel Hiltgen authored Sep 02, 2025
  
  0cc90a81
31 Aug, 2025 2 commits
- ml: fix struct field name in comment (#12123) · e42300f2
  pxwanglu authored Sep 01, 2025
  
  e42300f2
- readme: add NOMYO Router to community integrations (#12129) · 66e73809
  alpha-nerd-nomyo authored Aug 31, 2025
  
  66e73809
29 Aug, 2025 2 commits

perf: build graph for next batch async to keep GPU busy (#11863) · 517807cd

Daniel Hiltgen authored Aug 29, 2025

* perf: build graph for next batch in parallel to keep GPU busy

This refactors the main run loop of the ollama runner to perform the main GPU
intensive tasks (Compute+Floats) in a go routine so we can prepare the next
batch in parallel to reduce the amount of time the GPU stalls waiting for the
next batch of work.

* tests: tune integration tests for ollama engine

This tunes the integration tests to focus more on models supported
by the new engine.

517807cd

Always filter devices (#12108) · ead4a9a1

Daniel Hiltgen authored Aug 29, 2025

* Always filter devices

Avoid crashing on unsupported AMD iGPUs

* Remove cuda device filtering

This interferes with mixed setups

ead4a9a1

28 Aug, 2025 1 commit
- readme: add Neuro SAN to community integrations (#12109) · 4383a3ab
  ofrancon authored Aug 28, 2025
  
  4383a3ab
27 Aug, 2025 2 commits

ggml: Avoid allocating CUDA primary context on unused GPUs · 9d97e6a9

Jesse Gross authored Aug 26, 2025

The recent memory management changes caused all GPUs to be visible
to the runner, regardless of whether they are ultimately used. This
caused CUDA devices to allocate a primary context (~300 MB VRAM) on
each GPU, for each model. This is unnecessary, so we can both avoid
touching GPUs that we exclude in the early stage of allocation and
freeing the memory for any that we touch but don't use.

The issue will continue to exist for the old engine, since it touches
all devices during initialization.

9d97e6a9

fix keep alive (#12041) · 10815324
Michael Yang authored Aug 27, 2025

10815324

26 Aug, 2025 1 commit
- convert(gptoss): mxfp4 to ggml layout to avoid jit conversion (#12018) · 59412fbb
  Michael Yang authored Aug 26, 2025
```
* convert: return bytes written

* ggml flavor mxfp4

* simplify jit conversion

* comment
```
  59412fbb