Commits · 0a066cfd91abdddc6ee172776974a6720a3072d3 · OpenDAS / ollama

20 Jun, 2025 1 commit
- Reapply "feat: incremental gguf parser (#10822)" (#11114) (#11119) · 0a066cfd
  Michael Yang authored Jun 20, 2025
```
* Reapply "feat: incremental gguf parser (#10822)" (#11114)

This reverts commit a6e64fbd.

* fix older ggufs
```
  0a066cfd
18 Jun, 2025 2 commits
- Revert "feat: incremental gguf parser (#10822)" (#11114) · a6e64fbd
  Jeffrey Morgan authored Jun 18, 2025
```
This reverts commit 6b04cad7.
```
  a6e64fbd
- cache: fix comment function name in cache.go (#11110) · 60cfa2a2
  曹家巧 authored Jun 18, 2025
  
  60cfa2a2
12 Jun, 2025 2 commits

tools: loosen tool parsing to allow for more formats (#11030) · 9f8a18ec
Jeffrey Morgan authored Jun 12, 2025

9f8a18ec

feat: incremental gguf parser (#10822) · 6b04cad7

Michael Yang authored Jun 12, 2025



* incremental gguf parser
* gguf: update test to not rely on gguf on disc
* re-use existing create gguf
* read capabilities from gguf kv
* kv exists
* update tests
* s/doneFunc/successFunc/g
* new buffered reader

---------
Co-authored-by: Bruce MacDonald <brucewmacdonald@gmail.com>

6b04cad7

07 Jun, 2025 1 commit
- Revert "server: add model capabilities to the list endpoint (#10174)" (#11004) · 09d308d6
  Jeffrey Morgan authored Jun 06, 2025
```
This reverts commit 09430011.
```
  09d308d6
06 Jun, 2025 1 commit
- move thinking logic into its own package (#10990) · a3b6886b
  Devon Rifkin authored Jun 06, 2025
```
move thinking logic into its own package
```
  a3b6886b
05 Jun, 2025 1 commit
- export ThinkingParser · 0683efa6
  Devon Rifkin authored Jun 05, 2025
  
  0683efa6
04 Jun, 2025 1 commit
- server: add model capabilities to the list endpoint (#10174) · 09430011
  JasonHonKL authored Jun 05, 2025
  
  09430011
29 May, 2025 1 commit

add thinking support to the api and cli (#10584) · 5f57b0ef

Devon Rifkin authored May 28, 2025

- Both `/api/generate` and `/api/chat` now accept a `"think"`
  option that allows specifying whether thinking mode should be on or
  not
- Templates get passed this new option so, e.g., qwen3's template can
  put `/think` or `/no_think` in the system prompt depending on the
  value of the setting
- Models' thinking support is inferred by inspecting model templates.
  The prefix and suffix the parser uses to identify thinking support is
  also automatically inferred from templates
- Thinking control & parsing is opt-in via the API to prevent breaking
  existing API consumers. If the `"think"` option is not specified, the
  behavior is unchanged from previous versions of ollama
- Add parsing for thinking blocks in both streaming/non-streaming mode
  in both `/generate` and `/chat`
- Update the CLI to make use of these changes. Users can pass `--think`
  or `--think=false` to control thinking, or during an interactive
  session they can use the commands `/se...

5f57b0ef

27 May, 2025 1 commit
- server: abort download on empty digest · 9239a254
  Kyle Steere authored May 27, 2025
```
Signed-off-by: Kyle Steere <kyle.steere@chainguard.dev>
```
  9239a254
24 May, 2025 1 commit
- server: add hint to the error message when model path access fails (#10843) · eda472df
  frob authored May 24, 2025
  
  eda472df
23 May, 2025 1 commit
- tools: refactor tool call parsing and enable streaming (#10415) · e8b981fa
  Parth Sareen authored May 23, 2025
  
  e8b981fa
22 May, 2025 2 commits

sched: fix runner leak during reloading unload (#10819) · d950ff12

Daniel Hiltgen authored May 22, 2025

When the same model is being reloaded rapidly with client connections
being canceled before the model finishes loading, the queued unload
event could cause a leak of runners by deleting a different runner from
the loaded list.

d950ff12

server: improve tensor quantization fallback logic (#10806) · fbe6ae28

Bruce MacDonald authored May 22, 2025

Fall back to alternative quantization types when a tensor's dimensions aren't divisible by the block size required for the original desired quantization type. If retried quantization types fail, the system ultimately falls back to F16 (half-precision floating point) which has a block size of 1 and can handle any tensor dimension.

fbe6ae28

21 May, 2025 1 commit

remove support for multiple ggufs in a single file (#10722) · 61aeaf7e

Michael Yang authored May 21, 2025

* remove support for multiple ggufs in a single file

this was an attempt to make it easier to import multimodal models into
ollama. this was rarely used and error prone so remove it

* fix: create fused model from blob

61aeaf7e

19 May, 2025 2 commits

avoid kv truncation during create (#10761) · 1a0cfd08
Daniel Hiltgen authored May 19, 2025

1a0cfd08

ggml: Seperate tensor load from backend creation · 94ab428e

Jesse Gross authored Apr 17, 2025

Currently, when the backend is created, the tensors are loaded at the
same time, which is a slow operation. This separates them to be two
steps:
 - Create backend, including enumerating tensors and memory allocation
 - Loading tensor data

This allows more flexibility in managing model loading.

94ab428e

14 May, 2025 2 commits
- fix crash in old clients with quantization progress (#10710) · ff80718e
  Daniel Hiltgen authored May 14, 2025
```
Older clients assumed the digest was at least 19 characters long so increase the size
of the dummy digest to avoid array out of bounds crashes.
```
  ff80718e
- chore: update mllama to use ollama engine (#10637) · 23125648
  Michael Yang authored May 13, 2025
  
  23125648
13 May, 2025 1 commit
- server: add webp image input support (#10653) · c7f4ae7b
  Jeffrey Morgan authored May 12, 2025
  
  c7f4ae7b
12 May, 2025 3 commits

Follow up to #10363 (#10647) · 9d6df908

Daniel Hiltgen authored May 12, 2025

The quantization PR didn't block all unsupported file types,
which this PR fixes.  It also updates the API docs to reflect
the now reduced set of supported types.

9d6df908

convert: quantize from safetensors needs kv (#10675) · ad035ad5

Bruce MacDonald authored May 12, 2025

When creating a quantized model from safetensors we
need the array KV values to be loaded.Changing this
value to -1 loads the KV values on the returned
layer to be used and saved during quantization.

ad035ad5

feat: add trace log level (#10650) · f95a1f2b
Michael Yang authored May 12, 2025
```
reduce prompt log to trace level
```
f95a1f2b

08 May, 2025 2 commits

fix: stream accumulator exits early (#10593) · 0d6e35d3

Michael Yang authored May 08, 2025

the stream accumulator exits as soon as it sees `api.ProgressResponse(status="success")` which isn't strictly correctly
since some requests may have multiple successes, e.g. `/api/create` when the source model needs to be pulled.

0d6e35d3

lint: enable usetesting, disable tenv (#10594) · 6e9a7a25
Michael Yang authored May 08, 2025

6e9a7a25

07 May, 2025 2 commits

sched: fix race leading to orphaned runners (#10599) · 5e380c3b

Daniel Hiltgen authored May 07, 2025

If a model is loading, and the request context is canceled during the load
by a client closing the connection, and another request is inbound for the
same model with a different configuration (context size, etc.) thus requiring
a reload, two unload events can be in flight. The first shuts down the
original model load, but the second one caused the loss of the new
reloading runner reference, thus triggering the leak.

The primary fix is detecting the duplicate unload and ignoring the second
instance. The load routine is also hardened to ensure we detect
clobbering an already present runner and unload it with a warning.

5e380c3b

api: remove unused RetrieveModelResponse type (#10603) · 392de840
Jeffrey Morgan authored May 06, 2025

392de840

06 May, 2025 3 commits
- server: send 405 instead of 404 for unallowed methods (#10275) · 4090aca9
  Devon Rifkin authored May 06, 2025
```
Fixes: #5483
```
  4090aca9
- server: remove internal cmd (#10595) · 92ce438d
  Michael Yang authored May 06, 2025
  
  92ce438d
- Move quantization to new backend (#10363) · 42481045
  Daniel Hiltgen authored May 06, 2025
```
* Move quantization logic to GGML via new backend

This moves the model aware logic to Go code and calls GGMLs quantization code for model creation.

* Remove "add model quantizations"

This is no longer needed now that quantization is implemented in Go+GGML code directly.
```
  42481045
05 May, 2025 1 commit
- server: fix panic when runner.Options is nil (#10566) · 1703d147
  Jeffrey Morgan authored May 05, 2025
  
  1703d147
03 May, 2025 1 commit

sched: logging improvements (#10550) · 76ea735a

Daniel Hiltgen authored May 03, 2025

This enhances our logging in the scheduler. The initial "waiting for server" log
no longer claims an initial error state (now "not responding" which better reflects
the actual state). Runners now have slog wiring to report more details about the
runner, including PID.

76ea735a

01 May, 2025 1 commit
- image: add vision capability for projector-based models (#10509) · e6d2d041
  frob authored May 02, 2025
```
Co-authored-by: Richard Lyons <frob@cloudstaff.com>
```
  e6d2d041
30 Apr, 2025 2 commits

strip out thinking tags in message history for qwen3 & r1 (#10490) · ad3c7c9b

Devon Rifkin authored Apr 30, 2025

* strip out thinking tags in message history for qwen3 & r1

This is in advance of "proper" support where we'll make reasoning
configurable and we'll parse out thinking/reasoning tags and provide
them to the caller. These models expect there to be no thinking tags in
the message history, so this should improve quality

* parse model names instead of hacky prefix check

ad3c7c9b

Fix "Stopping..." scheduler hang (#10487) · 415c8fcc

Daniel Hiltgen authored Apr 30, 2025

* Adjust initial scheduler refCount

Ensure we only set the refCount on success

* sched: fix lock order inversion deadlock

Under certain race conditions, there was a scenario where the scheduler would
get into a deadlock while trying to update free space information while a model
was trying to unload.

415c8fcc

29 Apr, 2025 1 commit

lower default num parallel to 2 · fe5b9bb2

Devon Rifkin authored Apr 29, 2025

this is in part to "pay" for #10452, which doubled the default context length. The combination isn't fully neutral though, because even though the old 4x2k limit and the new 2x4k limit are memory equivalent, the 1x fallback is larger with 4k

fe5b9bb2

28 Apr, 2025 1 commit
- Revert "increase default context length to 4096 (#10364)" · dd93e1af
  Devon Rifkin authored Apr 28, 2025
```
This reverts commit 424f6486.
```
  dd93e1af
25 Apr, 2025 2 commits

explicitly decode maxarraysize 1024 · 340448d2
Michael Yang authored Apr 25, 2025

340448d2

fix superfluous call to WriteHeader · 214a7678

Michael Yang authored Apr 24, 2025

the first call to http.ResponseWriter.Write implicitly calls WriteHeader
with http.StatusOK if it hasn't already been called. once WriteHeader
has been called, subsequent calls has no effect. Write is called when
JSON encoding progressUpdateJSON{}. calls to
http.ResponseWriter.WriteHeader after the first encode is useless and
produces a warning:

http: superfluous response.WriteHeader call from github.com/ollama/ollama/server/internal/registry.(*statusCodeRecorder).WriteHeader (server.go:77)

214a7678