Commits · c68f367ef6688972de6798e631a7aa50c48af763 · OpenDAS / ollama

02 Oct, 2025 1 commit

Update GGML to b6646 (#12245) · c68f367e

Daniel Hiltgen authored Oct 02, 2025

Notable EOLs with this change:
- MacOS v12 and v13 are no longer supported (v14+ required)
- AMD gfx900 and gfx906 are no longer supported

c68f367e

01 Oct, 2025 1 commit

Use runners for GPU discovery (#12090) · bc8909fb

Daniel Hiltgen authored Oct 01, 2025

This revamps how we discover GPUs in the system by leveraging the Ollama
runner. This should eliminate inconsistency between our GPU discovery and the
runners capabilities at runtime, particularly for cases where we try to filter
out unsupported GPUs. Now the runner does that implicitly based on the actual
device list. In some cases free VRAM reporting can be unreliable which can
leaad to scheduling mistakes, so this also includes a patch to leverage more
reliable VRAM reporting libraries if available.

Automatic workarounds have been removed as only one GPU leveraged this, which
is now documented. This GPU will soon fall off the support matrix with the next
ROCm bump.

Additional cleanup of the scheduler and discovery packages can be done in the
future once we have switched on the new memory management code, and removed
support for the llama runner.

bc8909fb

22 Sep, 2025 2 commits
- docs: update cloud.md for cloud models · af060eb2
  jmorganca authored Sep 19, 2025
  
  af060eb2
- docs: move turbo.md to cloud.md · ae5c3300
  jmorganca authored Sep 19, 2025
  
  ae5c3300
15 Sep, 2025 1 commit
- doc: show how to clear the cgo cache (#12298) · 93c64ea1
  Daniel Hiltgen authored Sep 15, 2025
  
  93c64ea1
11 Sep, 2025 1 commit
- feat: add dimensions field to embed requests (#12242) · feb18cd7
  Michael Yang authored Sep 11, 2025
```
* feat: add field to truncate embeddings

* add openai embeddings for dimensions
```
  feb18cd7
10 Sep, 2025 1 commit

Add v12 + v13 cuda support (#12000) · 17a023f3

Daniel Hiltgen authored Sep 10, 2025

* Add support for upcoming NVIDIA Jetsons

The latest Jetsons with JetPack 7 are moving to an SBSA compatible model and
will not require building a JetPack specific variant.

* cuda: bring back dual versions

This adds back dual CUDA versions for our releases,
with v11 and v13 to cover a broad set of GPUs and
driver versions.

* win: break up native builds in build_windows.ps1

* v11 build working on windows and linux

* switch to cuda v12.8 not JIT

* Set CUDA compression to size

* enhance manual install linux docs

17a023f3

08 Sep, 2025 1 commit
- docs: show how to debug nvidia init failures (#12216) · 950d33aa
  Daniel Hiltgen authored Sep 08, 2025
```
This debug setting can help troubleshoot obscure initialization failures.
```
  950d33aa
15 Aug, 2025 1 commit
- docs: added missing comma in 'Ollama's Javascript library'' (#11915) · 883d0312
  Thomas Pelster authored Aug 15, 2025
  
  883d0312
14 Aug, 2025 1 commit
- doc: clarify both rocm and main bundle necessary (#11900) · 7ccfd97a
  Daniel Hiltgen authored Aug 14, 2025
```
Some users expect the rocm bundles to be self-sufficient, but are designed to be additive.
```
  7ccfd97a
06 Aug, 2025 3 commits
- docs: update the faq (#11760) · 44bc36d0
  Patrick Devine authored Aug 06, 2025
  
  44bc36d0
- Update downloading to pulling in api.md (#11170) · 8a75e9ee
  Gao feng authored Aug 07, 2025
```
update api.md to make it consist with code.
https://github.com/ollama/ollama/blob/main/server/download.go#L447
```
  8a75e9ee
- docs: update turbo model name (#11707) · 4742e12c
  Parth Sareen authored Aug 05, 2025
  
  4742e12c
05 Aug, 2025 1 commit
- docs: add docs for Ollama Turbo (#11687) · ee92ca3e
  Jeffrey Morgan authored Aug 05, 2025
  
  ee92ca3e
28 Jul, 2025 1 commit
- docs: fix typos and remove trailing whitespaces (#11554) · 3515cc37
  Yoshi authored Jul 28, 2025
  
  3515cc37
22 Jul, 2025 1 commit
- Update linux.md (#11462) · 4151ef8c
  ycomiti authored Jul 22, 2025
  
  4151ef8c
17 Jul, 2025 1 commit
- docs: add the no-Modelfile function of `ollama create` (#9077) · 802ad16c
  frob authored Jul 17, 2025
  
  802ad16c
16 Jul, 2025 1 commit
- docs: fix typo in macos.md (#11425) · 2e3fd86d
  Marcelo Fornet authored Jul 16, 2025
  
  2e3fd86d
11 Jul, 2025 1 commit
- docs: update modelfile.md to reflect current default num_ctx (#11189) · 4261a3b0
  先知 authored Jul 11, 2025
```
As in the commit 44b466ee, the default context length has been increased to 4096.
```
  4261a3b0
08 Jul, 2025 2 commits

doc: add MacOS docs (#11334) · 66fb8575
Daniel Hiltgen authored Jul 08, 2025
```
also removes stale model dir instructions for windows
```
66fb8575

Reduce default parallelism to 1 (#11330) · 20c3266e

Daniel Hiltgen authored Jul 08, 2025

The current scheduler algorithm of picking the paralellism based on available
VRAM complicates the upcoming dynamic layer memory allocation algorithm. This
changes the default to 1, with the intent going forward that parallelism is
explicit and will no longer be dynamically determined. Removal of the dynamic
logic will come in a follow up.

20c3266e

07 Jul, 2025 2 commits
- add `tool_name` to api.md (#11326) · 43107b15
  Parth Sareen authored Jul 07, 2025
  
  43107b15
- template: add tool result compatibility (#11294) · 1f91cb0c
  Parth Sareen authored Jul 07, 2025
  
  1f91cb0c
05 Jul, 2025 1 commit
- doc: add NVIDIA blackwell to supported list (#11307) · 9d60bb44
  Daniel Hiltgen authored Jul 05, 2025
  
  9d60bb44
23 Jun, 2025 1 commit

Re-remove cuda v11 (#10694) · 1c6669e6

Daniel Hiltgen authored Jun 23, 2025

* Re-remove cuda v11

Revert the revert - drop v11 support requiring drivers newer than Feb 23

This reverts commit c6bcdc42.

* Simplify layout

With only one version of the GPU libraries, we can simplify things down somewhat.  (Jetsons still require special handling)

* distinct sbsa variant for linux arm64

This avoids accidentally trying to load the sbsa cuda libraries on
a jetson system which results in crashes.

* temporary prevent rocm+cuda mixed loading

1c6669e6

18 Jun, 2025 1 commit
- benchmark: remove unused benchmark test (#11120) · 8bcb3125
  Jeffrey Morgan authored Jun 18, 2025
```
Removes a test under benchmark/ that is unused
```
  8bcb3125
07 Jun, 2025 2 commits
- docs: update link to AMD drivers in linux.md (#10973) · fc030961
  Krzysztof Jeziorny authored Jun 07, 2025
  
  fc030961
- Revert "server: add model capabilities to the list endpoint (#10174)" (#11004) · 09d308d6
  Jeffrey Morgan authored Jun 06, 2025
```
This reverts commit 09430011.
```
  09d308d6
06 Jun, 2025 1 commit
- docs: fix typo in development.md (#10998) · c6a6d729
  Hunter Wittenborn authored Jun 06, 2025
  
  c6a6d729
04 Jun, 2025 1 commit
- server: add model capabilities to the list endpoint (#10174) · 09430011
  JasonHonKL authored Jun 05, 2025
  
  09430011
29 May, 2025 1 commit

add thinking support to the api and cli (#10584) · 5f57b0ef

Devon Rifkin authored May 28, 2025

- Both `/api/generate` and `/api/chat` now accept a `"think"`
  option that allows specifying whether thinking mode should be on or
  not
- Templates get passed this new option so, e.g., qwen3's template can
  put `/think` or `/no_think` in the system prompt depending on the
  value of the setting
- Models' thinking support is inferred by inspecting model templates.
  The prefix and suffix the parser uses to identify thinking support is
  also automatically inferred from templates
- Thinking control & parsing is opt-in via the API to prevent breaking
  existing API consumers. If the `"think"` option is not specified, the
  behavior is unchanged from previous versions of ollama
- Add parsing for thinking blocks in both streaming/non-streaming mode
  in both `/generate` and `/chat`
- Update the CLI to make use of these changes. Users can pass `--think`
  or `--think=false` to control thinking, or during an interactive
  session they can use the commands `/se...

5f57b0ef

24 May, 2025 1 commit
- docs: remove unsupported quantizations (#10842) · 66238981
  frob authored May 24, 2025
  
  66238981
13 May, 2025 1 commit

Revert "remove cuda v11 (#10569)" (#10692) · c6bcdc42

Daniel Hiltgen authored May 13, 2025

Bring back v11 until we can better warn users that their driver
is too old.

This reverts commit fa393554.

c6bcdc42

12 May, 2025 1 commit

Follow up to #10363 (#10647) · 9d6df908

Daniel Hiltgen authored May 12, 2025

The quantization PR didn't block all unsupported file types,
which this PR fixes.  It also updates the API docs to reflect
the now reduced set of supported types.

9d6df908

08 May, 2025 1 commit
- api: remove unused sampling parameters (#10581) · fa9973cd
  Jeffrey Morgan authored May 08, 2025
  
  fa9973cd
07 May, 2025 1 commit

remove cuda v11 (#10569) · fa393554

Daniel Hiltgen authored May 06, 2025

This reduces the size of our Windows installer payloads by ~256M by dropping
support for nvidia drivers older than Feb 2023. Hardware support is unchanged.

Linux default bundle sizes are reduced by ~600M to 1G.

fa393554

05 May, 2025 1 commit

api: remove unused or unsupported api options (#10574) · 3b2d2c83

Jeffrey Morgan authored May 05, 2025

Some options listed in api/types.go are not supported in
newer models, or have been deprecated in the past. This is
the first of a series of PRs to clean up the API options

3b2d2c83

29 Apr, 2025 1 commit
- config: update default context length to 4096 · 44b466ee
  Devon Rifkin authored Apr 28, 2025
  
  44b466ee
28 Apr, 2025 1 commit
- Revert "increase default context length to 4096 (#10364)" · dd93e1af
  Devon Rifkin authored Apr 28, 2025
```
This reverts commit 424f6486.
```
  dd93e1af
22 Apr, 2025 1 commit

increase default context length to 4096 (#10364) · 424f6486

Devon Rifkin authored Apr 22, 2025

* increase default context length to 4096

We lower the default numParallel from 4 to 2 and use these "savings" to
double the default context length from 2048 to 4096.

We're memory neutral in cases when we previously would've used
numParallel == 4, but we add the following mitigation to handle some
cases where we would have previously fallen back to 1x2048 due to low
VRAM: we decide between 2048 and 4096 using a runtime check, choosing
2048 if we're on a one GPU system with total VRAM of <= 4 GB. We
purposefully don't check the available VRAM because we don't want the
context window size to change unexpectedly based on the available VRAM.

We plan on making the default even larger, but this is a relatively
low-risk change we can make to quickly double it.

* fix tests

add an explicit context length so they don't get truncated. The code
that converts -1 from being a signal for doing a runtime check isn't
running as part of these tests.

* tweak small gpu message

* clarify context length default

also make it actually show up in `ollama serve --help`

424f6486