Commits · 45c47393748f739b2c58ed2597ca1ce0baecff88 · OpenDAS / ollama

16 Dec, 2025 1 commit

types: ConfigV2 and RootFS (#13504) · 45c47393

Bruce MacDonald authored Dec 16, 2025

Refactored the ConfigV2 and RootFS types from server/images.go to a new types/model/config.go file under the model package. Updated all references to use model.ConfigV2 and model.RootFS. This allows for use in other projects without worrying about compiling the c code in the llama package.

45c47393

11 Dec, 2025 2 commits
- openai: add v1/responses support (#13351) · 1eb5e759
  Devon Rifkin authored Dec 11, 2025
```
Only supporting the stateless part of the API.

Doc updates to come once this is shipped.

Closes: #9659
```
  1eb5e759
- routes: add logprobs in tool calls (#13238) · 1c4e85b4
  EasonLin authored Dec 11, 2025
  
  1c4e85b4
08 Dec, 2025 1 commit

truncation: fixed runner truncation logic + removed server truncation (#12839) · e082d60a

nicole pardal authored Dec 08, 2025

This PR consolidates all embedding prompt-length checking, truncation, and prompt token counting into the runner to ensure a single source of truth.

e082d60a

05 Dec, 2025 1 commit

fix(api): correct Content-Type header for /api/chat and /api/generate when... · 31b8c6a2

Sos Pogosyan authored Dec 05, 2025


fix(api): correct Content-Type header for /api/chat and /api/generate when using cloud models (#13279)

---------
Co-authored-by: Pogosyan Sos <sos_pogosyan@MacBook-Pro-Sos.local>
Co-authored-by: Patrick Devine <patrick@infrahq.com>

31b8c6a2

20 Nov, 2025 1 commit
- Parser for Cogito v2 (#13145) · d70e9355
  Grace authored Nov 19, 2025
  
  d70e9355
11 Nov, 2025 1 commit

server: add logprobs and top_logprobs support to Ollama's API (#12899) · 59241c5b

Baptiste Jamin authored Nov 11, 2025



Adds logprobs support to Ollama's API including support for Ollama's
OpenAI-compatible API. By specifying the new 'logprobs' boolean parameter
in the API, Ollama will return the log probabilities for each token generated.
'top_logprobs', an integer value can also be specified up to the value 20.
When specified, the API will also provide the number of most likely tokens to
return at each token position
Co-authored-by: Baptiste Jamin <baptiste@crisp.chat>

59241c5b

05 Nov, 2025 1 commit

Add Tool Call ID (#12956) · 809b9c68

Grace authored Nov 04, 2025



* routes/types: add tool call id

---------
Co-authored-by: ParthSareen <parth.sareen@ollama.com>

809b9c68

29 Oct, 2025 1 commit
- feat(model): add qwen3vl (#12665) · 7d25b9e1
  Michael Yang authored Oct 28, 2025
  
  7d25b9e1
28 Oct, 2025 1 commit
- Revert "server: Consolidate embedding truncation in runner (#12730)" (#12810) · 29f63f37
  Patrick Devine authored Oct 28, 2025
```
This reverts commit 5d347f6d.
```
  29f63f37
27 Oct, 2025 1 commit

server: Consolidate embedding truncation in runner (#12730) · 5d347f6d

nicole pardal authored Oct 27, 2025

Currently, checking the length of prompts for embeddings to ensure
they fit in the context window (and possible truncation) occurs in
two places - the Ollama server and runner. This can lead to
inconsistencies in both the checks and reported number of tokens
processed. Since we have to do this processing in the runner, this
consolidates all of the logic there.

5d347f6d

25 Oct, 2025 1 commit
- cloud: set the proxy content-type to the same as local models (#12759) · b97eb2b8
  Patrick Devine authored Oct 25, 2025
  
  b97eb2b8
22 Oct, 2025 1 commit
- cloud: don't error sending empty messages (#12724) · d515aed6
  Patrick Devine authored Oct 21, 2025
  
  d515aed6
16 Oct, 2025 1 commit

renderers: add global flag for setting [img] tags (#12669) · 65fb3ff4

Jeffrey Morgan authored Oct 16, 2025

Adds a temporary global flag to renderers that causes renderers to always
render images as [img]. In a follow up change, we will consider making this
the default, and this flag could eventually be removed

65fb3ff4

11 Oct, 2025 2 commits

Reapply "add truncate and shift parameters" (#12582) · 6544e147
Jeffrey Morgan authored Oct 11, 2025

6544e147

routes: fix built-in renderers for `api/generate` · 6db8da99

Devon Rifkin authored Oct 11, 2025

Made it so when api/generate builds up a message array and generates the
prompt it now goes through the same function as `api/chat` for
consistency. This is where we hook the optional built-in renderers to
bypass templates, which was missing for `api/generate` before this
change.

Closes: #12578

6db8da99

10 Oct, 2025 1 commit
- thinking: allow `"think": false` for non-thinking models (#12555) · d681cd7c
  Patrick Devine authored Oct 09, 2025
  
  d681cd7c
09 Oct, 2025 3 commits
- routes: structured outputs for gpt-oss (#12460) · 77060d46
  Parth Sareen authored Oct 08, 2025
  
  77060d46
- Revert "add truncate and shift parameters (#12519)" (#12545) · 7d965258
  Jeffrey Morgan authored Oct 08, 2025
```
This reverts commit 6a62b894.
```
  7d965258
- add truncate and shift parameters (#12519) · 6a62b894
  Jeffrey Morgan authored Oct 08, 2025
  
  6a62b894
08 Oct, 2025 1 commit
- thinking: turn on thinking mode for all reasoning models (#12533) · 90d429f5
  Patrick Devine authored Oct 08, 2025
  
  90d429f5
05 Oct, 2025 1 commit

openai: refactor to split compat layer and middleware · 2c2f4dea

Devon Rifkin authored Oct 05, 2025

This makes the core openai compat layer independent of the middleware
that adapts it to our particular gin routes

2c2f4dea

01 Oct, 2025 2 commits

Use runners for GPU discovery (#12090) · bc8909fb

Daniel Hiltgen authored Oct 01, 2025

This revamps how we discover GPUs in the system by leveraging the Ollama
runner. This should eliminate inconsistency between our GPU discovery and the
runners capabilities at runtime, particularly for cases where we try to filter
out unsupported GPUs. Now the runner does that implicitly based on the actual
device list. In some cases free VRAM reporting can be unreliable which can
leaad to scheduling mistakes, so this also includes a patch to leverage more
reliable VRAM reporting libraries if available.

Automatic workarounds have been removed as only one GPU leveraged this, which
is now documented. This GPU will soon fall off the support matrix with the next
ROCm bump.

Additional cleanup of the scheduler and discovery packages can be done in the
future once we have switched on the new memory management code, and removed
support for the llama runner.

bc8909fb

fix keep alive · 35ac4eb1

Michael Yang authored Sep 30, 2025

this reference to keep alive was missed in #12041 so chat has a
diffferent behaviour than generate

35ac4eb1

23 Sep, 2025 1 commit

auth: fix problems with the ollama keypairs (#12373) · 64883e3c

Patrick Devine authored Sep 22, 2025

* auth: fix problems with the ollama keypairs

This change adds several fixes including:
  - reading in the pubkey files correctly
  - fixing the push unit test to create a keypair file in a temp directory
  - not return 500 errors for normal status error

64883e3c

18 Sep, 2025 3 commits

server: add unauthorized error to remote chat handler (#12338) · 22ccdd74
Jeffrey Morgan authored Sep 18, 2025

22ccdd74

harmony: remove special casing in routes.go · e7f56ef3

Devon Rifkin authored Sep 18, 2025

Now that we have a built-in parser abstraction, which was introduced in
<https://github.com/ollama/ollama/pull/12248>, we can modify our harmony
parser to match this and then get rid of nearly all of the
harmony-specific logic in routes.go. We do have a small amount of
code that turns the parser on by default if the architecture matches and
no other built-in parser was provided.

The built-in parser interface was modified in order to handle harmony's
prefill and tool name translation requirements.

e7f56ef3

fix(integration): check truncated length (#12337) · ceac416e
Michael Yang authored Sep 18, 2025

ceac416e

17 Sep, 2025 2 commits
- server: skip parsing initial <think> if provided in the prompt for /api/generate (#12289) · 9b8187b4
  frob authored Sep 18, 2025
  
  9b8187b4
- engine: add remote proxy (#12307) · 8b894933
  Patrick Devine authored Sep 17, 2025
  
  8b894933
15 Sep, 2025 3 commits

model: implement bert in ollama engine (#9080) · 3f6642f6

Michael Yang authored Sep 15, 2025

* fix truncate

* s/SentencePieceModel/SentencePiece/

* bert

* wordpiece

* refactor pooling

* more tokenizers

* normalize embeddings

3f6642f6

address comments · 472feec2
Devon Rifkin authored Sep 15, 2025

472feec2

add qwen3-coder tool support · 47991940

Devon Rifkin authored Sep 11, 2025

The format qwen3-coder uses is relatively unique, both in rendering and
in parsing. To implement parsing, I wrote a custom parser in similar
style to harmony. For the rendering, I found that the logic would be
much more difficult to follow in a template, so I introduced the concept
of a built-in renderer that uses go code, rather than a template to
generate prompts.

I set us up for future built-in parsers and renderers by making it so
they can be specified in a Modelfile like so:

```
RENDERER "qwen3-coder"
PARSER "qwen3-coder"
```

These need to be provided explicitly because the architecture alone is
not enough to understand what format the model expects to receive, and
what format we expect it to output (e.g., qwen3-coder is `qwen3moe`,
which includes other qwen3-family models as well)

I haven't converted harmony to be one of these "built-ins" yet, since
some of it is in flux with the changes @ParthSareen has been making to
move harmony to the runner. It is likely that many other built-ins will
need to move to the runner as well, but I'm able to slightly defer that
decision since qwen3-coder doesn't have thinking (and therefore doesn't
need to be in the runner to make structured outputs work). I expect to
unify harmony with this approach very soon.

Whether a particular model supports tools or thinking was previously
inferred from templates, but without a template we now also use the
parser itself to declare what it supports. If we have future models that
re-use the same parsing format, but have different capabilities, we'll
want to parameterize them and give them different names to be specified
as a `PARSER`.

Misc changes:

- I worked on the renderer by diffing outputs from the reference
  implementation and ours. To make it easier to do this, I extended
  <https://github.com/ollama/ollama/pull/11875> to also support
  returning the prompt via the openai compat layer

47991940

12 Sep, 2025 2 commits
- Revert "runner: move harmony to runner (#12052)" · 92b96d54
  jmorganca authored Sep 12, 2025
```
This reverts commit 1a558f98.
```
  92b96d54
- Revert "runner: simplify parser entrypoints in runner (#12233)" · 9d56e63d
  jmorganca authored Sep 12, 2025
```
This reverts commit 8d6fffae.
```
  9d56e63d
11 Sep, 2025 1 commit
- feat: add dimensions field to embed requests (#12242) · feb18cd7
  Michael Yang authored Sep 11, 2025
```
* feat: add field to truncate embeddings

* add openai embeddings for dimensions
```
  feb18cd7
10 Sep, 2025 1 commit
- runner: simplify parser entrypoints in runner (#12233) · 8d6fffae
  Parth Sareen authored Sep 10, 2025
  
  8d6fffae
08 Sep, 2025 1 commit
- runner: move harmony to runner (#12052) · 1a558f98
  Parth Sareen authored Sep 08, 2025
  
  1a558f98
27 Aug, 2025 1 commit
- fix keep alive (#12041) · 10815324
  Michael Yang authored Aug 27, 2025
  
  10815324
22 Aug, 2025 1 commit
- server: skip parsing initial <think> if provided in the prompt (#12024) · 4be4dc87
  Jeffrey Morgan authored Aug 22, 2025
  
  4be4dc87