Commits · 1e6a28bf5b1fbcf540bb4fc4e2f89408ebd04384 · OpenDAS / ollama

28 Apr, 2024 1 commit

Daniel Hiltgen authored Apr 28, 2024

Prior refactoring passes accidentally removed the logic to bypass VRAM
checks for CPU loads. This adds that back, along with test coverage.

This also fixes loaded map access in the unit test to be behind the mutex which was
likely the cause of various flakes in the tests.

d6e3b645

25 Apr, 2024 1 commit
- Reload model if `num_gpu` changes (#3920) · 00b0699c
  Jeffrey Morgan authored Apr 25, 2024
```
* reload model if `num_gpu` changes

* dont reload on -1

* fix tests
```
  00b0699c
24 Apr, 2024 3 commits
- Restructure loading conditional chain · 36a6dacc
  Bryce Reitano authored Apr 24, 2024
  
  36a6dacc
- Provide variable ggml for TestLoad · ceb0e26e
  Bryce Reitano authored Apr 24, 2024
  
  ceb0e26e
- Move ggml loading to when we attempt fitting · 284e02be
  Bryce Reitano authored Apr 24, 2024
  
  284e02be
23 Apr, 2024 2 commits

Harden sched TestLoad · d8851cb7
Daniel Hiltgen authored Apr 23, 2024
```
Give the go routine a moment to deliver the expired event
```
d8851cb7

Request and model concurrency · 34b9db5a

Daniel Hiltgen authored Mar 30, 2024

This change adds support for multiple concurrent requests, as well as
loading multiple models by spawning multiple runners. The default
settings are currently set at 1 concurrent request per model and only 1
loaded model at a time, but these can be adjusted by setting
OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS.

34b9db5a