Commits · 4e320b8b90b8a698fc3c057a3f54cbabe59b543a · OpenDAS / ollama

14 Mar, 2025 2 commits

server/internal/chunks: remove chunks package (#9755) · 4e320b8b
Blake Mizerany authored Mar 14, 2025

4e320b8b

server/internal/client: use chunksums for concurrent blob verification (#9746) · eb2b22b0

Blake Mizerany authored Mar 13, 2025

Replace large-chunk blob downloads with parallel small-chunk
verification to solve timeout and performance issues. Registry users
experienced progressively slowing download speeds as large-chunk
transfers aged, often timing out completely.

The previous approach downloaded blobs in a few large chunks but
required a separate, single-threaded pass to read the entire blob back
from disk for verification after download completion.

This change uses the new chunksums API to fetch many smaller
chunk+digest pairs, allowing concurrent downloads and immediate
verification as each chunk arrives. Chunks are written directly to their
final positions, eliminating the entire separate verification pass.

The result is more reliable downloads that maintain speed throughout the
transfer process and significantly faster overall completion, especially
over unstable connections or with large blobs.

eb2b22b0

13 Mar, 2025 17 commits
- Merge pull request #9703 from ollama/mxyng/gemma3-memory · 4ea4d2b1
  Michael Yang authored Mar 13, 2025
```
count gemma3 vision tensors
```
  4ea4d2b1
- count non-repeating vision layers · 8d76fa23
  Michael Yang authored Mar 13, 2025
  
  8d76fa23
- docs: Add OLLAMA_ORIGINS for browser extension support (#9643) · 74b44fdf
  Bradley Erickson authored Mar 13, 2025
  
  74b44fdf
- fix divide by zero · 65b88c54
  Michael Yang authored Mar 13, 2025
  
  65b88c54
- roughly count gemma3 graph · a422ba39
  Michael Yang authored Mar 13, 2025
```
the largest operation is by far (q @ k) so just count that for
simplicity
```
  a422ba39
- count all vision tensors · d2ec2237
  Michael Yang authored Mar 12, 2025
  
  d2ec2237
- count gemma3 vision tensors · 033cec23
  Michael Yang authored Mar 12, 2025
  
  033cec23
- Merge pull request #9741 from ollama/mxyng/visionless · 543240fb
  Michael Yang authored Mar 13, 2025
```
fix: error if image requested without vision model
```
  543240fb
- add verbose mode to the show command (#9640) · 4bed7392
  Patrick Devine authored Mar 13, 2025
```
Add metadata and tensor information to the show command to be able to
see more information about a model. This outputs the same data as
shown on the model details page on ollama.com
```
  4bed7392
- fix: change default context size for gemma3 (#9744) · 80c7ce38
  Patrick Devine authored Mar 13, 2025
  
  80c7ce38
- Merge pull request #9742 from ollama/mxyng/engine-error-embeddings · ccfd41c4
  Michael Yang authored Mar 13, 2025
```
fix: error on models that don't support embeddings
```
  ccfd41c4
- Update model/model.go · 3e102b7d
  Michael Yang authored Mar 13, 2025
```
Co-authored-by: Jeffrey Morgan <jmorganca@gmail.com>
```
  3e102b7d
- engine: error on embeddings; not currently implemented · ec46f328
  Michael Yang authored Mar 13, 2025
  
  ec46f328
- fix: error if image requested without vision model · 5e2e0b46
  Michael Yang authored Mar 13, 2025
  
  5e2e0b46
- Merge pull request #9688 from Shane-XB-Qian/debug_mistype_lld · 45a13b1d
  Michael Yang authored Mar 13, 2025
```
ollama-debug.c: correct mistype
```
  45a13b1d
- sample: separate softmax and temperature transforms (#9732) · 5c0b6639
  Parth Sareen authored Mar 13, 2025
  
  5c0b6639
- ollama-debug.c: change 'ld' to 'PRIi64' · 30d7a59b
  shane.xb.qian authored Mar 13, 2025
```
* macOS has different definition per info from @mxyng
```
  30d7a59b
12 Mar, 2025 8 commits
- sample: do all sorting in topK · 4aeb67ef
  ParthSareen authored Mar 12, 2025
  
  4aeb67ef
- sample: simplify top_k=0 sorting · 3ba91634
  ParthSareen authored Mar 12, 2025
  
  3ba91634
- sample: use container/heap for top_k · 1b7433b7
  ParthSareen authored Mar 12, 2025
  
  1b7433b7
- models/gemma3: remove final logit softcap (#9692) · a70820da
  Bruce MacDonald authored Mar 12, 2025
```
Softcap isn't in the whitepaper/implementation for the language model so we should remove it. There is no discernible difference in output with it removed.
```
  a70820da
- cli: adding support ctrl-n/p like general cli (#9136) · 6b45b1d6
  Shane-XB-Qian authored Mar 12, 2025
```
Signed-off-by: shane.xb.qian <shane.qian@foxmail.com>
```
  6b45b1d6
- ollama-debug.c: correct mistype · 85ab5520
  shane.xb.qian authored Mar 12, 2025
```
Signed-off-by: shane.xb.qian <shane.qian@foxmail.com>
```
  85ab5520
- cli: don't exit for invalid model during /load. (#9576) · b3af953a
  frob authored Mar 12, 2025
```
Co-authored-by: Richard Lyons <frob@cloudstaff.com>
```
  b3af953a
- Adding Gemma 3 to readme (#9671) · ad4e0bf3
  Michael authored Mar 12, 2025
  
  ad4e0bf3
11 Mar, 2025 13 commits
- Merge pull request #9661 from ollama/gemma · aee28501
  Michael Yang authored Mar 11, 2025
```
engine: add gemma support
```
  aee28501
- all: address linter errors · 83f0ec82
  jmorganca authored Mar 11, 2025
  
  83f0ec82
- kvcache: fix tests by adding AvgPool2D stub · c6b6938b
  jmorganca authored Mar 11, 2025
  
  c6b6938b
- model: add more spm tokenizer tests · fb4664fc
  jmorganca authored Mar 11, 2025
  
  fb4664fc
- model: validate left and right pairs before merging them · 20e35938
  jmorganca authored Mar 11, 2025
  
  20e35938
- use 2d pooling · 63a39406
  Michael Yang authored Mar 11, 2025
  
  63a39406
- llm: auto detect models that require Ollama Engine (#1 ) · ab39e08e
  Daniel Hiltgen authored Mar 11, 2025
  
  ab39e08e
- add trailing \n\n after <end_of_image> to match reference implementation · 11bfa627
  jmorganca authored Mar 11, 2025
  
  11bfa627
- reduce kernel size, add TODO for loading from config · f63e62e5
  jmorganca authored Mar 11, 2025
  
  f63e62e5
- Revert "Allow models to force a new batch" · 65b0f329
  jmorganca authored Mar 11, 2025
```
This reverts commit c7eae586b899083acebcd9b3847b89ea78c2850c.
```
  65b0f329
- Allow models to force a new batch · 06007c0a
  Jesse Gross authored Mar 10, 2025
```
This is useful for a few things:
 - Work around bugs, such as having 2 images in one batch
 - Keep the image in a single batch for fully connected attention
 - Improve performance by not evaluating embeddings multiple times
```
  06007c0a
- Disable causal attention based on batch index · a8e83a76
  Jesse Gross authored Mar 10, 2025
```
Currently we are using positions, which are relative to a
sequence and may not be unique.
```
  a8e83a76
- Restrict Gemma to a single image per request · 47500550
  Jesse Gross authored Mar 10, 2025
  
  47500550