Commits · f6c29409dc0823fcb0b42e8138ea3e208d6b5edf · OpenDAS / ollama

29 Oct, 2025 2 commits
- docs: add new cloud model + fix openai redirect (#12812) · f6c29409
  Jeffrey Morgan authored Oct 28, 2025
  
  f6c29409
- feat(model): add qwen3vl (#12665) · 7d25b9e1
  Michael Yang authored Oct 28, 2025
  
  7d25b9e1
28 Oct, 2025 12 commits
- embed: add distance correlation test for library embed models (#12796) · 36d64fb5
  Patrick Devine authored Oct 28, 2025
  
  36d64fb5
- docs: update readme and links (#12809) · d828517e
  Parth Sareen authored Oct 28, 2025
  
  d828517e
- Fix vulkan PCI ID and ID handling (#12775) · 14977a93
  Daniel Hiltgen authored Oct 28, 2025
```
* Fix vulkan PCI ID and ID handling

Intel GPUs may not report PCI IDs which was leading to incorrect overlap
detection.  Switch to using the existing PCI IDs, however AMD GPUs claim not to
report PCI IDs, but actually do, so try anyway, as this is required for ADLX to
find the GPUs on Windows. Numeric IDs lead to scheduling problems, so this also
switches Vulkan to use UUID based IDs. The GPU discovery patches have been
squashed into a single patch to simplify future rebases.

* review comments
```
  14977a93
- Revert "server: Consolidate embedding truncation in runner (#12730)" (#12810) · 29f63f37
  Patrick Devine authored Oct 28, 2025
```
This reverts commit 5d347f6d.
```
  29f63f37
- docs: add docs for docs.ollama.com (#12805) · 3d99d977
  Parth Sareen authored Oct 28, 2025
  
  3d99d977
- docs: rename to mdx to setup docs site (#12804) · 6d02a43a
  Parth Sareen authored Oct 28, 2025
  
  6d02a43a
- Revert "docs: add reference to docs.ollama.com (#12800)" (#12803) · 5483497d
  Parth Sareen authored Oct 28, 2025
```
This reverts commit 934dd9e1.
```
  5483497d
- docs: add reference to docs.ollama.com (#12800) · 934dd9e1
  Parth Sareen authored Oct 28, 2025
  
  934dd9e1
- s/From*Slice/From*s/ (#12255) · 1188f408
  Michael Yang authored Oct 28, 2025
  
  1188f408
- embedding tests: added check against exact base64 string (#12790) · 15c7d30d
  nicole pardal authored Oct 28, 2025
  
  15c7d30d
- Merge pull request #12793 from ollama/drifkin/12792_renderer-parser-from · 98623171
  Devon Rifkin authored Oct 28, 2025
```
create: inherit FROM model's renderer/parser
```
  98623171
- gemma3: make embedding non-causal (#12297) · ec9eb28f
  Michael Yang authored Oct 27, 2025
  
  ec9eb28f
27 Oct, 2025 2 commits

create: inherit FROM model's renderer/parser · 1bdd8169

Devon Rifkin authored Oct 27, 2025

On main, the `RENDERER` and `PARSER` fields from the `Modelfile` don't
get propagated to a new model created with a `req.From` parameter. This
is easily triggered via `ollama run qwen3-coder`, then running some save
command like `/save qwen3-coder-custom`.

Added a regression test for this, and then open the config for the
"from" model in order to use its renderer/parser as a default for the
new model. This will fix the CLI and also API-based creates.

Fixes: https://github.com/ollama/ollama/issues/12792

1bdd8169

server: Consolidate embedding truncation in runner (#12730) · 5d347f6d

nicole pardal authored Oct 27, 2025

Currently, checking the length of prompts for embeddings to ensure
they fit in the context window (and possible truncation) occurs in
two places - the Ollama server and runner. This can lead to
inconsistencies in both the checks and reported number of tokens
processed. Since we have to do this processing in the runner, this
consolidates all of the logic there.

5d347f6d

25 Oct, 2025 1 commit
- cloud: set the proxy content-type to the same as local models (#12759) · b97eb2b8
  Patrick Devine authored Oct 25, 2025
  
  b97eb2b8
23 Oct, 2025 4 commits

llm: Change memory allocation backoff from exponential to incremental · ad6f6a1d

Jesse Gross authored Oct 23, 2025

If we create a memory layout that should fit based on report free VRAM
but allocation still fails, we start applying a backoff. This reduces
free VRAM by an exponential percentage (1%, 2%, 4%...). However, the
points chosen tend to be too dense at the beginning and too sparse at
the end. Therefore, this switches to an incremental backoff (10%, 20%,
30%...).

ad6f6a1d

readme: add VT Code project to terminal community integrations (#12749) · 6723a40b
Vinh Nguyen authored Oct 24, 2025

6723a40b

DRY out the runner lifecycle code (#12540) · 3258a89b

Daniel Hiltgen authored Oct 23, 2025

* DRY out the runner lifecycle code

Now that discovery uses the runners as well, this unifies the runner spawning code
into a single place.  This also unifies GPU discovery types with the newer ml.DeviceInfo

* win: make incremental builds better

Place build artifacts in discrete directories so incremental builds don't have to start fresh

* Adjust sort order to consider iGPUs

* handle cpu inference oom scenarios

* review comments

3258a89b

kvcache: Remove special case for reservation mask · 1c093e97

Jesse Gross authored Oct 22, 2025

We currently short circuit generation of the cache mask and just
generate an empty tensor of the correct size. However, in some
cases, this can also skip a cast operation. This can result in the
worst case graph being not fully worst case.

We don't actually need the fast path for mask generation, so it's
better to just use the normal code path.

1c093e97

22 Oct, 2025 4 commits

llamarunner: Record the time for all batches during prompt processing · a8d9c264

Jesse Gross authored Oct 16, 2025

Currently, we only record the time for the last batch when processing
the prompt. This results in unrealistically high numbers for the
old llama runner.

Before:
total duration:       31.273112939s
load duration:        4.97054657s
prompt eval count:    32768 token(s)
prompt eval duration: 235.137439ms
prompt eval rate:     139356.80 tokens/s
eval count:           1873 token(s)
eval duration:        18.173182374s
eval rate:            103.06 tokens/s

After:
total duration:       30.024798033s
load duration:        4.758588663s
prompt eval count:    32768 token(s)
prompt eval duration: 7.779621548s
prompt eval rate:     4212.03 tokens/s
eval count:           1769 token(s)
eval duration:        17.148014223s
eval rate:            103.16 tokens/s

a8d9c264

tools: parse tool calls that don't conform to ("name": name, "arguments": args} (#12738) · 0334e67f
frob authored Oct 22, 2025

0334e67f
embeddings: base64 encoding fix (#12715) · e0ead1ad
nicole pardal authored Oct 22, 2025

e0ead1ad
cloud: don't error sending empty messages (#12724) · d515aed6
Patrick Devine authored Oct 21, 2025

d515aed6

20 Oct, 2025 5 commits
- runner: always truncate embeddings requests (#12714) · 5fe7ba1b
  Jeffrey Morgan authored Oct 20, 2025
  
  5fe7ba1b
- fs(ggml): fill in arch prefix if necessary (#12646) · d2b63c19
  Michael Yang authored Oct 20, 2025
  
  d2b63c19
- model/parsers: remove warning for missing <think> tag for qwen3-vl (#12713) · 94f110b3
  Jeffrey Morgan authored Oct 20, 2025
  
  94f110b3
- cuda: get driver version after props (#12707) · 5d22953b
  Daniel Hiltgen authored Oct 20, 2025
```
Users on Windows without GPUs are reporting errors relating to
cudaDriverGetVersion with the device set to -1.  This ensures we only grab the
driver once we're enumerating actual devices.
```
  5d22953b
- rocm: give it more time to bootstrap (#12681) · d245dffe
  Daniel Hiltgen authored Oct 20, 2025
```
Some users are hitting timeouts.  We'd like to make this faster, but for now make sure we don't timeout too aggressively.
```
  d245dffe
18 Oct, 2025 2 commits

contiguous input per layer (#12686) · bc1a818f
Daniel Hiltgen authored Oct 17, 2025
```
Co-authored-by: Michael Yang <git@mxy.ng>
```
bc1a818f

win: more verbose load failures (#12683) · ba2253dc

Daniel Hiltgen authored Oct 17, 2025

When loading the dynamic libraries, if something goes wrong report some
details.  Unfortunately this wont explain which dependencies are missing,
but this breadcrumb in the logs should help us diagnose GPU discovery
failures.

ba2253dc

17 Oct, 2025 1 commit

test: harden scheduler tests (#12662) · 68e04c7f

Daniel Hiltgen authored Oct 17, 2025

* test: harden scheduler tests

This removes reschedDelay which was stale code, and adds
a new configurable timeout for the waitForVRAMRecovery so
tests can now set the timeout to be very short to avoid the
scheduler getting stuck and hitting a test timeout.

* test: tune tests for partial loads

Give stress tests more time when the model is split between CPU/GPU

68e04c7f

16 Oct, 2025 7 commits

cuda: tidy up CC settings (#12668) · 27067993
Daniel Hiltgen authored Oct 16, 2025
```
8.7 is Jetpack only, so no need on x86 builds
10.3 covers [G]B300
```
27067993

renderers: add global flag for setting [img] tags (#12669) · 65fb3ff4

Jeffrey Morgan authored Oct 16, 2025

Adds a temporary global flag to renderers that causes renderers to always
render images as [img]. In a follow up change, we will consider making this
the default, and this flag could eventually be removed

65fb3ff4

Grace/qwen3 thinking (#12647) · e2a0b244

Grace authored Oct 16, 2025

* changing initial status to take into consideration prefill

* Add seperate strings for content and thinking builder

* thinking tests

* remove white space from string before closing think tag

e2a0b244

cuda: bring back CC 5.2 (#12666) · 1813ff85

Daniel Hiltgen authored Oct 16, 2025

Forward compat on the newer driver doesn't seem to be working.
This should get 5.2 working on newer drivers again.

1813ff85

test: add a few missing embedding models (#12661) · b531777a
Daniel Hiltgen authored Oct 16, 2025

b531777a
Revert "Workaround broken NVIDIA iGPU free VRAM data (#12490)" (#12642) · fe3ec8db
Daniel Hiltgen authored Oct 16, 2025
```
The workaround has been moved into the underlying C++ code.

This reverts commit e4340667.
```
fe3ec8db
vulkan: Get FilterID from Backend for Vulkan (#12655) · c7441342
Thomas Stocker authored Oct 16, 2025
```
* vulkan: Get FilterID from Backend for Vulkan

* Fixing patch
```
c7441342