Commits · 8852220f59c12cf3165f5643d38453ffecbb722d · OpenDAS / ollama

18 Dec, 2025 3 commits
- add REQUIRES command to Modelfile (#13361) · 8852220f
  Jeffrey Morgan authored Dec 18, 2025
  
  8852220f
- Revert "Omit args and params in tool function def and calls (#13516)" (#13518) · 522c11a7
  Grace authored Dec 17, 2025
```
This reverts commit 0fadeffa.
```
  522c11a7
- Omit args and params in tool function def and calls (#13516) · 0fadeffa
  Grace authored Dec 17, 2025
  
  0fadeffa
17 Dec, 2025 1 commit
- types: add nested property support for tool definitions (#13508) · 1c094038
  Parth Sareen authored Dec 17, 2025
  
  1c094038
12 Dec, 2025 1 commit
- docs: fix link to modelfile.mdx (#13220) · 93d45d7a
  Alexander Gusak authored Dec 12, 2025
  
  93d45d7a
04 Dec, 2025 1 commit
- ci: restore previous linter rules (#13322) · 854d40ed
  Jeffrey Morgan authored Dec 03, 2025
  
  854d40ed
01 Dec, 2025 1 commit

api/client: handle non-json streaming errors (#13007) · 5b6a8e60

Bruce MacDonald authored Dec 01, 2025

While processing the response stream during a chat or generation if an error is occurred it is parsed and returned to the user. The issue with the existing code is that this assumed the response would be valid JSON, which is not a safe assumption and caused cryptic error messages to be displayed due to parsing failures:
`invalid character 'i' looking for beginning of value`

This change updates the stream function to return the raw error string if it cant be parsed as JSON. This should help with debugging issues by making sure the actual error reaches the user.

5b6a8e60

13 Nov, 2025 1 commit
- logprob: add bytes to logprobs (#13068) · c1149875
  Parth Sareen authored Nov 13, 2025
  
  c1149875
11 Nov, 2025 1 commit

server: add logprobs and top_logprobs support to Ollama's API (#12899) · 59241c5b

Baptiste Jamin authored Nov 11, 2025



Adds logprobs support to Ollama's API including support for Ollama's
OpenAI-compatible API. By specifying the new 'logprobs' boolean parameter
in the API, Ollama will return the log probabilities for each token generated.
'top_logprobs', an integer value can also be specified up to the value 20.
When specified, the API will also provide the number of most likely tokens to
return at each token position
Co-authored-by: Baptiste Jamin <baptiste@crisp.chat>

59241c5b

06 Nov, 2025 1 commit
- api: add omitempty to required tool function parameter type (#12989) · 30fcc719
  Jeffrey Morgan authored Nov 06, 2025
  
  30fcc719
05 Nov, 2025 1 commit

Add Tool Call ID (#12956) · 809b9c68

Grace authored Nov 04, 2025



* routes/types: add tool call id

---------
Co-authored-by: ParthSareen <parth.sareen@ollama.com>

809b9c68

15 Oct, 2025 1 commit
- types: send index for tool calls (#12625) · c4c5a4a0
  Parth Sareen authored Oct 14, 2025
  
  c4c5a4a0
13 Oct, 2025 1 commit

Qwen3VL Cloud Parser and Renderer (#12526) · 05982a95

Grace authored Oct 13, 2025



* working (other than tool call is the incorrect order) for tool calls and tools

* Tests work, other than image tags (tests do not go through server) and tools (not in the correct order, but contents are the same)

* testing for qwen3vl parser - toolparser is working

* made changes to JSON tool parser, wraps the TollCallFunction with a TollCall object

* Working parser for thinking models - assumes state of thinking, emits unambiguous content in thinking, does not call tool call in thinking

* changed the parser to start with collecting content

* thinking prefill

* add hasThinkingSupport parameter to parser

* qwen3-vl -> qwen3-vl-instruct for renderer/parser

* Add hasThinkingSupport=false to QwenVLParser

---------
Co-authored-by: Devon Rifkin <drifkin@drifkin.net>

05982a95

11 Oct, 2025 1 commit
- Reapply "add truncate and shift parameters" (#12582) · 6544e147
  Jeffrey Morgan authored Oct 11, 2025
  
  6544e147
09 Oct, 2025 2 commits
- Revert "add truncate and shift parameters (#12519)" (#12545) · 7d965258
  Jeffrey Morgan authored Oct 08, 2025
```
This reverts commit 6a62b894.
```
  7d965258
- add truncate and shift parameters (#12519) · 6a62b894
  Jeffrey Morgan authored Oct 08, 2025
  
  6a62b894
08 Oct, 2025 1 commit
- thinking: turn on thinking mode for all reasoning models (#12533) · 90d429f5
  Patrick Devine authored Oct 08, 2025
  
  90d429f5
23 Sep, 2025 1 commit

auth: fix problems with the ollama keypairs (#12373) · 64883e3c

Patrick Devine authored Sep 22, 2025

* auth: fix problems with the ollama keypairs

This change adds several fixes including:
  - reading in the pubkey files correctly
  - fixing the push unit test to create a keypair file in a temp directory
  - not return 500 errors for normal status error

64883e3c

17 Sep, 2025 1 commit
- engine: add remote proxy (#12307) · 8b894933
  Patrick Devine authored Sep 17, 2025
  
  8b894933
15 Sep, 2025 1 commit

add qwen3-coder tool support · 47991940

Devon Rifkin authored Sep 11, 2025

The format qwen3-coder uses is relatively unique, both in rendering and
in parsing. To implement parsing, I wrote a custom parser in similar
style to harmony. For the rendering, I found that the logic would be
much more difficult to follow in a template, so I introduced the concept
of a built-in renderer that uses go code, rather than a template to
generate prompts.

I set us up for future built-in parsers and renderers by making it so
they can be specified in a Modelfile like so:

```
RENDERER "qwen3-coder"
PARSER "qwen3-coder"
```

These need to be provided explicitly because the architecture alone is
not enough to understand what format the model expects to receive, and
what format we expect it to output (e.g., qwen3-coder is `qwen3moe`,
which includes other qwen3-family models as well)

I haven't converted harmony to be one of these "built-ins" yet, since
some of it is in flux with the changes @ParthSareen has been making to
move harmony to the runner. It is likely that many other built-ins will
need to move to the runner as well, but I'm able to slightly defer that
decision since qwen3-coder doesn't have thinking (and therefore doesn't
need to be in the runner to make structured outputs work). I expect to
unify harmony with this approach very soon.

Whether a particular model supports tools or thinking was previously
inferred from templates, but without a template we now also use the
parser itself to declare what it supports. If we have future models that
re-use the same parsing format, but have different capabilities, we'll
want to parameterize them and give them different names to be specified
as a `PARSER`.

Misc changes:

- I worked on the renderer by diffing outputs from the reference
  implementation and ours. To make it easier to do this, I extended
  <https://github.com/ollama/ollama/pull/11875> to also support
  returning the prompt via the openai compat layer

47991940

11 Sep, 2025 1 commit
- feat: add dimensions field to embed requests (#12242) · feb18cd7
  Michael Yang authored Sep 11, 2025
```
* feat: add field to truncate embeddings

* add openai embeddings for dimensions
```
  feb18cd7
27 Aug, 2025 1 commit
- fix keep alive (#12041) · 10815324
  Michael Yang authored Aug 27, 2025
  
  10815324
22 Aug, 2025 1 commit
- api: implement stringer for ToolFunctionParameters (#12038) · d3450dd5
  Jeffrey Morgan authored Aug 22, 2025
  
  d3450dd5
15 Aug, 2025 1 commit
- server: add debug option for printing out prompt instead of calling model · 8de1da47
  Devon Rifkin authored Aug 15, 2025
  
  8de1da47
12 Aug, 2025 1 commit
- fix(openai): handle reasoning_effort (#11868) · d0cf6c82
  Michael Yang authored Aug 12, 2025
  
  d0cf6c82
05 Aug, 2025 2 commits

tools: support anyOf types · 30f8a68c

Devon Rifkin authored Aug 05, 2025

afaik gpt-oss is the first model that meaningfully transforms tool
function definitions in its template. We found that relatively common
definitions that include `anyOf` were not working because the template
was assuming that types were always defined via a `type` field.

anyOf allows for fully recursive types, so I exposed a
`toTypeScriptType()` function to handle this recursive logic in go and
keep the templates cleaner. The gpt-oss templates will need to be
updated to use this.

We should keep building out our function definition support to more
fully support the parts of json schema that make sense for this use
case, but in the meantime this will unblock some users (e.g., zed's
ollama integration w/ gpt-oss). Probably the most urgent is proper array
support

30f8a68c

gpt-oss (#11672) · fa7776fd

Michael Yang authored Aug 05, 2025



* bf16

* tests

* gpt-oss

* enable gptoss for engine

* rough estimate

* convert to mxfp4

* handle safetensors U8

* clamp glu/linear

* update tokenizer

* MXFP4 support

This implements the Open Compute Microscaling (MX) FP4 format
as a tensor type with backend implementations focusing
on mulmat and mulmatid on CPU, CUDA, and Metal.

* Unit tests for MXFP4 support

This exercises various operations and shapes on both CPU and GPU (if detected
on the system)

* cuda graph

* unit test adjustments

* cuda: optimize memory access

Read 4 bytes at a time (8 elements) when performing mul_mat_vec_mxfp4

* mac: fix crash on old macos versions

cblas_sgemm is only supported on v13.3 and up, however bf16 is
only supported on v14+ so we were falling back to ggml-blas and
crashing on bf16 tensors.  Checking for the function being null
seems to be the simplest way to condittionally avoid registering the
backend.

* server: Minimum context length for gptoss

This model requires a minimum context length of 8192 to function
effectively. Users can set higher values through all normal mechanisms
but lower values will be silently reset.

* ggml: Multiply by numParallel for gptoss sliding window

When computing the graph size estimate, the context size is already
multiplied by numParallel so estimates reflect that. However, since
sliding window models use a smaller, fixed context size, they need
to manually take numParallel into account.

* gpt-oss integration

includes harmony parser and thinking levels, etc.

* fix sync

* fix tests

* fix lint

---------
Co-authored-by: Daniel Hiltgen <daniel@ollama.com>
Co-authored-by: Jesse Gross <jesse@ollama.com>
Co-authored-by: Devon Rifkin <drifkin@drifkin.net>

fa7776fd

16 Jul, 2025 1 commit

api: fix unreachable status err (#11423) · 92c2e8a5

Bruce MacDonald authored Jul 16, 2025

StatusError was unreachable, the client always checked for error messages in the response body first, and the server always includes error messages with HTTP error status codes.

92c2e8a5

08 Jul, 2025 1 commit

API/CLI context enhancements (#11331) · 34088dbc

Daniel Hiltgen authored Jul 08, 2025

* API: expose context size of loaded models

* CLI: add context UX

This adds a column in the ps output to show the models context size.

34088dbc

07 Jul, 2025 1 commit
- template: add tool result compatibility (#11294) · 1f91cb0c
  Parth Sareen authored Jul 07, 2025
  
  1f91cb0c
07 Jun, 2025 1 commit
- Revert "server: add model capabilities to the list endpoint (#10174)" (#11004) · 09d308d6
  Jeffrey Morgan authored Jun 06, 2025
```
This reverts commit 09430011.
```
  09d308d6
04 Jun, 2025 1 commit
- server: add model capabilities to the list endpoint (#10174) · 09430011
  JasonHonKL authored Jun 05, 2025
  
  09430011
29 May, 2025 1 commit

add thinking support to the api and cli (#10584) · 5f57b0ef

Devon Rifkin authored May 28, 2025

- Both `/api/generate` and `/api/chat` now accept a `"think"`
  option that allows specifying whether thinking mode should be on or
  not
- Templates get passed this new option so, e.g., qwen3's template can
  put `/think` or `/no_think` in the system prompt depending on the
  value of the setting
- Models' thinking support is inferred by inspecting model templates.
  The prefix and suffix the parser uses to identify thinking support is
  also automatically inferred from templates
- Thinking control & parsing is opt-in via the API to prevent breaking
  existing API consumers. If the `"think"` option is not specified, the
  behavior is unchanged from previous versions of ollama
- Add parsing for thinking blocks in both streaming/non-streaming mode
  in both `/generate` and `/chat`
- Update the CLI to make use of these changes. Users can pass `--think`
  or `--think=false` to control thinking, or during an interactive
  session they can use the commands `/se...

5f57b0ef

27 May, 2025 1 commit
- client: add request signing to the client (#10881) · aa25aff1
  Patrick Devine authored May 27, 2025
```
If OLLAMA_AUTH is set, sign each request w/ a timestamp and pass the signature in the token header
```
  aa25aff1
08 May, 2025 2 commits
- lint: enable usetesting, disable tenv (#10594) · 6e9a7a25
  Michael Yang authored May 08, 2025
  
  6e9a7a25
- api: remove unused sampling parameters (#10581) · fa9973cd
  Jeffrey Morgan authored May 08, 2025
  
  fa9973cd
07 May, 2025 1 commit
- api: remove unused RetrieveModelResponse type (#10603) · 392de840
  Jeffrey Morgan authored May 06, 2025
  
  392de840
05 May, 2025 1 commit

api: remove unused or unsupported api options (#10574) · 3b2d2c83

Jeffrey Morgan authored May 05, 2025

Some options listed in api/types.go are not supported in
newer models, or have been deprecated in the past. This is
the first of a series of PRs to clean up the API options

3b2d2c83

24 Apr, 2025 1 commit
- api: fix ImageData struct comment to expect raw image bytes (#10386) · 40b10eee
  Adrien Duermael authored Apr 23, 2025
  
  40b10eee
10 Apr, 2025 1 commit
- types: include the 'items' and '$defs' fields to properly handle "array" types (#10091) · ef65174d
  Tom Sheffler authored Apr 09, 2025
```
---------
Co-authored-by: Parth Sareen <parth.sareen@ollama.com>
```
  ef65174d