Commits · 3518aaef3318b47c63d3df9ef3bdd96dff7541ae · OpenDAS / ollama

"vscode:/vscode.git/clone" did not exist on "c627506ff35030eb1f64d4e926b7e5a177718b6e"

21 Jun, 2024 2 commits

Disable concurrency for AMD + Windows · 9929751c

Daniel Hiltgen authored Jun 19, 2024

Until ROCm v6.2 ships, we wont be able to get accurate free memory
reporting on windows, which makes automatic concurrency too risky.
Users can still opt-in but will need to pay attention to model sizes otherwise they may thrash/page VRAM or cause OOM crashes.
All other platforms and GPUs have accurate VRAM reporting wired
up now, so we can turn on concurrency by default.

9929751c

Enable concurrency by default · 17b7186c

Daniel Hiltgen authored May 06, 2024

This adjusts our default settings to enable multiple models and parallel
requests to a single model. Users can still override these by the same
env var settings as before. Parallel has a direct impact on
num_ctx, which in turn can have a significant impact on small VRAM GPUs
so this change also refines the algorithm so that when parallel is not
explicitly set by the user, we try to find a reasonable default that fits
the model on their GPU(s). As before, multiple models will only load
concurrently if they fully fit in VRAM.

17b7186c

19 Jun, 2024 2 commits
- Revert "Revert "gpu: add env var for detecting Intel oneapi gpus (#5076)"" · d34d88e4
  Daniel Hiltgen authored Jun 19, 2024
```
This reverts commit 755b4e4f.
```
  d34d88e4
- Revert "gpu: add env var for detecting Intel oneapi gpus (#5076)" · 755b4e4f
  Wang,Zhe authored Jun 19, 2024
```
This reverts commit 163cd3e7.
```
  755b4e4f
17 Jun, 2024 1 commit
- gpu: add env var for detecting Intel oneapi gpus (#5076) · 163cd3e7
  Jeffrey Morgan authored Jun 16, 2024
```
* gpu: add env var for detecting intel oneapi gpus

* fix build error
```
  163cd3e7
14 Jun, 2024 2 commits

Centralize GPU configuration vars · 6be309e1

Daniel Hiltgen authored May 08, 2024

This should aid in troubleshooting by capturing and reporting the GPU
settings at startup in the logs along with all the other server settings.

6be309e1

Support forced spreading for multi GPU · 5e8ff556

Daniel Hiltgen authored May 08, 2024

Our default behavior today is to try to fit into a single GPU if possible.
Some users would prefer the old behavior of always spreading across
multiple GPUs even if the model can fit into one.  This exposes that
tunable behavior.

5e8ff556

13 Jun, 2024 1 commit
- add OLLAMA_MODELS to envconfig (#5029) · 94618b23
  Patrick Devine authored Jun 13, 2024
  
  94618b23
12 Jun, 2024 1 commit
- move OLLAMA_HOST to envconfig (#5009) · c69bc19e
  Patrick Devine authored Jun 12, 2024
  
  c69bc19e
06 Jun, 2024 1 commit
- API app/browser access (#4879) · 1a29e9a8
  royjhan authored Jun 06, 2024
```
* API app/browser access

* Add tauri (resolves #2291, #4791, #3799, #4388)
```
  1a29e9a8
04 Jun, 2024 2 commits
- some gocritic · c895a7d1
  Michael Yang authored May 21, 2024
  
  c895a7d1
- nosprintfhostport · dad7a987
  Michael Yang authored May 21, 2024
  
  dad7a987
30 May, 2024 1 commit

Fix OLLAMA_LLM_LIBRARY with wrong map name and add more env vars to help message (#4663) · a03be181

Lei Jitang authored May 31, 2024



* envconfig/config.go: Fix wrong description of OLLAMA_LLM_LIBRARY
Signed-off-by: Lei Jitang <leijitang@outlook.com>

* serve: Add more env to help message of ollama serve

Add more enviroment variables to `ollama serve --help`
to let users know what can be configurated.
Signed-off-by: Lei Jitang <leijitang@outlook.com>

---------
Signed-off-by: Lei Jitang <leijitang@outlook.com>

a03be181

24 May, 2024 1 commit
- Move envconfig and consolidate env vars (#4608) · 4cc3be30
  Patrick Devine authored May 24, 2024
  
  4cc3be30