Commits · 05a43e078a89247dcc71c703c1bee2af97c1655d · OpenDAS / ollama

02 Oct, 2025 1 commit
- fix panic on bootstrapDevices (#12475) · 05a43e07
  Daniel Hiltgen authored Oct 01, 2025
```
Wrong index variable was used.
```
  05a43e07
01 Oct, 2025 3 commits

Use runners for GPU discovery (#12090) · bc8909fb

Daniel Hiltgen authored Oct 01, 2025

This revamps how we discover GPUs in the system by leveraging the Ollama
runner. This should eliminate inconsistency between our GPU discovery and the
runners capabilities at runtime, particularly for cases where we try to filter
out unsupported GPUs. Now the runner does that implicitly based on the actual
device list. In some cases free VRAM reporting can be unreliable which can
leaad to scheduling mistakes, so this also includes a patch to leverage more
reliable VRAM reporting libraries if available.

Automatic workarounds have been removed as only one GPU leveraged this, which
is now documented. This GPU will soon fall off the support matrix with the next
ROCm bump.

Additional cleanup of the scheduler and discovery packages can be done in the
future once we have switched on the new memory management code, and removed
support for the llama runner.

bc8909fb

Merge pull request #12461 from ollama/drifkin/qwen3-coder-tweaks · 6b50f2b9
Devon Rifkin authored Sep 30, 2025
```
qwen3-coder: fix tool definition type rendering
```
6b50f2b9

fix keep alive · 35ac4eb1

Michael Yang authored Sep 30, 2025

this reference to keep alive was missed in #12041 so chat has a
diffferent behaviour than generate

35ac4eb1

30 Sep, 2025 5 commits

ggml: Preallocate CUDA pool memory · 3d0b1734

Jesse Gross authored Sep 09, 2025

The GGML CUDA backend allocates additional memory for intermediate
results during calculation. This memory isn't currently allocated
during worst case graph reservation and therefore not included in
scheduling. This means that as these buffers potentially grow
with context length, we could crash.

This extends the memory allocation system down layer from the GGML
graph to the CUDA layer, preallocating the worst case memory there
as well.

Fixes #11753

3d0b1734

ggml: Backport scale kernel fixes · efaee8c2

Jesse Gross authored Sep 23, 2025

The GGML scale kernel uses signed 32-bit ints to represent
the number of elements in the tensor. For large images,
mistral-small3.2 overflows this, triggering CUDA errors due
to negative arguments.

Currently, this can happen when the user passes a large image
to mistral-small3.2. However, with upcoming changes to reserve
CUDA memory, it happens every time mistral-small is loaded as
we reserve using a worst case batch.

This patch is part of an upstream GGML commit and should be removed
after GGML is updated past 0a1b398 "ggml: add ops for WAN video model
(cuda && cpu) (#15669)".

Fixes #10388

efaee8c2

ggml: Remove allocation status reporting · 734b57da

Jesse Gross authored Sep 22, 2025

For each memory allocation we report the size of the (attempted)
allocation and whether it succeeded or failed. The latter status
reporting proved to be not that useful in practice as systems
such as Windows can automatically overflow from VRAM into RAM,
resultings in successful allocations even when there isn't
enough memory where we wanted.

As a result, this information is only used for debug logging,
which isn't worthwhile enough for the amount of code. It
also isn't fully accurate, as multiple allocations may result
in partial failures.

734b57da

qwen3-coder: fix tool definition type rendering · 83021fcf
Devon Rifkin authored Sep 30, 2025

83021fcf
build: call find_package to instantiate library paths · 0469861d
Michael Yang authored Sep 30, 2025

0469861d

26 Sep, 2025 2 commits

fix: correct condition for AMDGPU_TARGETS filtering logic (#12412) · c47154c0
羊撅撅 authored Sep 27, 2025

c47154c0

bugfix: restore the current runOptions if loading fails in the CLI (#12402) · b04e46da

Patrick Devine authored Sep 25, 2025

There are two bugs when using `/load <model>` for a model that doesn't exist, namely:
1. it will not restore the current model settings if the current model is a thinking model; and
2. it will crash is the current model is a non-thinking model

This bug fix saves the current runOptions and then restores them if the model load
doesn't happen. It also fixes the crash happening for non-thinking models.

b04e46da

25 Sep, 2025 4 commits

Merge pull request #12417 from ollama/drifkin/qwen3-coder-unicode · 34efbbd3
Devon Rifkin authored Sep 25, 2025
```
parsers: fix unicode handling for qwen3-coder
```
34efbbd3

parsers: fix unicode handling for qwen3-coder · 05ba4ca1

Devon Rifkin authored Sep 25, 2025

When trimming whitespace at the end of every chunk, we were iterating
backwards over the string byte-by-byte instead of rune-by-rune.

As an example of how this can cause corruption, suppose we have the
multi-byte character ✅ (`"\u2705"`), which is represented in utf-8 as
the three bytes `0xE2 0x9C 0x85`. It happens that `0x85` is NEL, which
passes `unicode.IsSpace()`. Because we were iterating byte-by-byte, this
caused us to mistakenly slice in the middle of the rune, removing `0x85`
and leaving `0xE2 0x9C`, which beyond being the incorrect place to
slice, is not even a valid utf-8 character.

`trailingWhitespaceLen()` was modified to count from the end in a
rune-aware way. Tests with various multibyte unicode characters were
also added.


Fixes: #12414

05ba4ca1

cli: add device signin flow when doing ollama push (#12405) · 5a56ff3c
Patrick Devine authored Sep 25, 2025

5a56ff3c
tools: handle the case where a tool call sends "arguments" or "parameters" as... · 2fba04b5
Gabe Goodhart authored Sep 25, 2025
```
tools: handle the case where a tool call sends "arguments" or "parameters" as a serialized json string (#12413)
```
2fba04b5

24 Sep, 2025 5 commits

Grace/deepseek v3 migration (#12385) · fbd82ba5

Grace authored Sep 24, 2025



* init deepseek model file

* temp removal of flash attention implementation

* shapes and proper, can make a pass

* query, key, value have good cosine similarity, but the max diff is a bit high

* Attention block is working! ** with eager for now, have not added the mask line

* Attention block is working! ** with eager for now, have not added the mask line

* working MoE at around 0.95 cosine sim

* added cosine similarity function

* Starting end to end structure

* Trying (and failing) to get rope to work, going to test full thing on tater

* running on tater36... just not the right outputs

* we have the right values for rope... but its still not working?

* chnage Extrapolation Factor to 1

* removed adding residuals twice, removed normalization from shared expert, refactored Norms (Attention, MLP) to be outside the (Attention, MLP) blocks and in the Transformer block instead, add cache setLayer

* Temporary modelfiles for cpu

* change kpass intermediate step to kv, two layer outputs [0,1] look fine

* this calls for 16 chicken nuggets

* whoops

* cleaning up code

* delete stuff we dont need

* getting rid of debug statements for llama cpp

* working with long contexts

* fix long context view error

* reverting some changes I made for files that are not apart of pr

* Added proper tokenizer for deeepseek3

* clean up model and go test

* remove Modelfile

* not passing the tests

* whoops

* how to pass the ci tests

* resolving some of the comments

* rename

* linted and renamed deepseek3 -> deepseek2

* remove name go

* addressed changes - main change was adopting qwen3 naming scheme

* I cannot with linters

* clean up logs

* clean up logs

---------
Co-authored-by: Grace Guo <graceguo@Graces-MBP.localdomain>
Co-authored-by: Grace Guo <graceguo@Graces-MacBook-Pro.local>
Co-authored-by: graceguo <graceguo@tater36.localdomain>

fbd82ba5

prefer ollama engine for qwen3moe (#12374) · 2e742544
Michael Yang authored Sep 24, 2025

2e742544
Merge pull request #12393 from ollama/drifkin/fix-built-ins · bbb195a6
Devon Rifkin authored Sep 23, 2025
```
harmony: don't sanitize built-ins
```
bbb195a6

harmony: don't sanitize built-ins · fd88cd7c

Devon Rifkin authored Sep 23, 2025

In #11910 we started sanitizing function names, but we accidentally were
modifying built-ins like `browser.open` to `browser_open`. This was
removing the special prompt rendering for built-ins, but this wasn't
immediately apparent since the models seem to be reasonably good at
remembering the built-ins even when presented with these slightly
renamed version. This fix prevents built-ins from ever being renamed.

fd88cd7c

fix: leaf alt name (#12390) · e1979c57

Michael Yang authored Sep 23, 2025

a leaf node with an alternative name gets all its alternatives names
added into the same branch rather than creating branches themselves

e1979c57

23 Sep, 2025 3 commits
- add pre:, suf: to tags (#12274) · bf78ed6e
  Michael Yang authored Sep 23, 2025
  
  bf78ed6e
- multi-regexp pretokenizer (#12325) · a40d427b
  Michael Yang authored Sep 23, 2025
  
  a40d427b
- auth: fix problems with the ollama keypairs (#12373) · 64883e3c
  Patrick Devine authored Sep 22, 2025
```
* auth: fix problems with the ollama keypairs

This change adds several fixes including:
  - reading in the pubkey files correctly
  - fixing the push unit test to create a keypair file in a temp directory
  - not return 500 errors for normal status error
```
  64883e3c
22 Sep, 2025 4 commits
- Merge pull request #12339 from ollama/drifkin/harmony-refactor-to-builtin · 41efdd40
  Devon Rifkin authored Sep 22, 2025
```
harmony: remove special casing in routes.go
```
  41efdd40
- tests: add single threaded history test (#12295) · c23e6f4c
  Daniel Hiltgen authored Sep 22, 2025
```
* tests: add single threaded history test

Also tidies up some existing tests to handle more model output variation

* test: add support for testing specific architectures
```
  c23e6f4c
- docs: update cloud.md for cloud models · af060eb2
  jmorganca authored Sep 19, 2025
  
  af060eb2
- docs: move turbo.md to cloud.md · ae5c3300
  jmorganca authored Sep 19, 2025
  
  ae5c3300
20 Sep, 2025 2 commits

Merge pull request #12358 from ollama/drifkin/qwen3-coder-ampersands · 3677842f
Devon Rifkin authored Sep 20, 2025
```
parsers: fix `&`s in qwen3coder parameter values
```
3677842f

parsers: fix `&`s in qwen3coder parameter values · 242df70a

Devon Rifkin authored Sep 20, 2025

In <https://github.com/ollama/ollama/issues/12357> we that the model
will output tool calls such as

```
<function=shell>
<parameter=command>
pwd && ls -la
</parameter>
</function>
```

We parse this using the approach of transforming into valid xml and then
using an xml parser. While we do transform the function and parameter
names, we weren't escaping the parameter values (which in this example
are invalid since `pwd && ls -la` contains unescaped ampersands).

This has been fixed by first transforming the tags in the same way, and
then walking the transformed string and escaping the text in between the
tags. This also fixes a case where `<` in the middle of a parameter
value would cause an xml parse failure.

Fixes: #12357

242df70a

19 Sep, 2025 1 commit
- gemma: fix rope scaling for qat models (#12348) · dba39b2e
  Patrick Devine authored Sep 19, 2025
```
* gemma: fix rope scaling for qat models

* gofumpt yourself
```
  dba39b2e
18 Sep, 2025 8 commits

fix: model load for unsupported embedding models (#12311) · 9f3a37fd

Michael Yang authored Sep 18, 2025

with #12181, there's now support for embeddings in ollama engine.
this is done by mutating the architecture and adding _embed when it
detects an embedding model. however this introduced a bug where if
an embedding model was run based on an existing ollama engine model
without an embedding implementation, e.g. llama4, it will pass the
initial arch support check but fail when actually loaded.

there's currently two entrypoints to creating a model. previously this
second entrypoint was necessary because calling model.New would also
load the model. since #11818, this is no longer th case so merge them
to reduce complexity

9f3a37fd

feat: qwen3 embed (#12301) · 7460259e
Michael Yang authored Sep 18, 2025
```
* cleanup

* use pooling.TypeNone

* pooling test

* qwen3 embed
```
7460259e
server: add unauthorized error to remote chat handler (#12338) · 22ccdd74
Jeffrey Morgan authored Sep 18, 2025

22ccdd74

build: avoid unbounded parallel builds (#12319) · 0c3d0e75

Daniel Hiltgen authored Sep 18, 2025

With the addition of cuda v13, on a clean setup, the level of parallelism
was causing docker desktop to become overwhelmed and compilers
were crashing. This limits to 8 parallel per build stage, with the ability
to override if you have many more cores available.

0c3d0e75

harmony: remove special casing in routes.go · e7f56ef3

Devon Rifkin authored Sep 18, 2025

Now that we have a built-in parser abstraction, which was introduced in
<https://github.com/ollama/ollama/pull/12248>, we can modify our harmony
parser to match this and then get rid of nearly all of the
harmony-specific logic in routes.go. We do have a small amount of
code that turns the parser on by default if the architecture matches and
no other built-in parser was provided.

The built-in parser interface was modified in order to handle harmony's
prefill and tool name translation requirements.

e7f56ef3

auth: check the permissions on the private key to see if it's readable (#12336) · eb0a5d44
Patrick Devine authored Sep 18, 2025

eb0a5d44
fix(integration): check truncated length (#12337) · ceac416e
Michael Yang authored Sep 18, 2025

ceac416e

convert: convert bf16 vision weights to fp16 (#12324) · 2717dce6

Patrick Devine authored Sep 17, 2025

This change moves back to converting bf16 vision weights to fp16,
specifically if they start with the name "v." (such as v.blk.0.attn_k.weight).

This fixes a bug where converted images are failing because they are trying
to call `im2col` which doesn't have a bf16 kernel in ggml.

2717dce6

17 Sep, 2025 2 commits
- server: skip parsing initial <think> if provided in the prompt for /api/generate (#12289) · 9b8187b4
  frob authored Sep 18, 2025
  
  9b8187b4
- engine: add remote proxy (#12307) · 8b894933
  Patrick Devine authored Sep 17, 2025
  
  8b894933