Commits · cb42e607c5cf4d439ad4d5a93ed13c7d6a09fc34 · OpenDAS / ollama

25 Jun, 2024 1 commit

llm: speed up gguf decoding by a lot (#5246) · cb42e607

Blake Mizerany authored Jun 24, 2024

Previously, some costly things were causing the loading of GGUF files
and their metadata and tensor information to be VERY slow:

  * Too many allocations when decoding strings
  * Hitting disk for each read of each key and value, resulting in a
    not-okay amount of syscalls/disk I/O.

The show API is now down to 33ms from 800ms+ for llama3 on a macbook pro
m3.

This commit also prevents collecting large arrays of values when
decoding GGUFs (if desired). When such keys are encountered, their
values are null, and are encoded as such in JSON.

Also, this fixes a broken test that was not encoding valid GGUF.

cb42e607

14 Jun, 2024 6 commits
- review comments and coverage · 6f351bf5
  Daniel Hiltgen authored Jun 05, 2024
  
  6f351bf5
- Prevent multiple concurrent loads on the same gpus · ff4f0cbd
  Daniel Hiltgen authored Jun 04, 2024
```
While models are loading, the VRAM metrics are dynamic, so try
to load on a GPU that doesn't have a model actively loading, or wait
to avoid races that lead to OOMs
```
  ff4f0cbd
- Refine CPU load behavior with system memory visibility · fc37c192
  Daniel Hiltgen authored Jun 03, 2024
  
  fc37c192
- Reintroduce nvidia nvml library for windows · 434dfe30
  Daniel Hiltgen authored Jun 03, 2024
```
This library will give us the most reliable free VRAM reporting on windows
to enable concurrent model scheduling.
```
  434dfe30
- Harden unload for empty runners · 48702dd1
  Daniel Hiltgen authored May 30, 2024
  
  48702dd1
- Support forced spreading for multi GPU · 5e8ff556
  Daniel Hiltgen authored May 08, 2024
```
Our default behavior today is to try to fit into a single GPU if possible.
Some users would prefer the old behavior of always spreading across
multiple GPUs even if the model can fit into one.  This exposes that
tunable behavior.
```
  5e8ff556
04 Jun, 2024 3 commits
- lint · e40145a3
  Michael Yang authored May 21, 2024
  
  e40145a3
- some gocritic · c895a7d1
  Michael Yang authored May 21, 2024
  
  c895a7d1
- replace x/exp/slices with slices · 04f3c12b
  Michael Yang authored May 21, 2024
  
  04f3c12b
24 May, 2024 1 commit
- Move envconfig and consolidate env vars (#4608) · 4cc3be30
  Patrick Devine authored May 24, 2024
  
  4cc3be30
21 May, 2024 1 commit

Correct typo in error message (#4535) · 4434d7f4

Sang Park authored May 22, 2024

The spelling of the term "request" has been corrected, which was previously mistakenly written as "requeset" in the error log message.

4434d7f4

14 May, 2024 2 commits
- Remove VRAM convergence check for windows · ec231a79
  Daniel Hiltgen authored May 14, 2024
```
The APIs we query are optimistic on free space, and windows pages
VRAM, so we don't have to wait to see reported usage recover on unload
```
  ec231a79
- Ollama `ps` command for showing currently loaded models (#4327) · 68459888
  Patrick Devine authored May 13, 2024
  
  68459888
10 May, 2024 2 commits
- Always use the sorted list of GPUs · 4142c3ef
  Daniel Hiltgen authored May 10, 2024
```
Make sure the first GPU has the most free space
```
  4142c3ef
- Don't clamp ctx size in `PredictServerFit` (#4317) · bb6fd022
  Jeffrey Morgan authored May 10, 2024
```
* dont clamp ctx size in `PredictServerFit`

* minimum 4 context

* remove context warning
```
  bb6fd022
09 May, 2024 1 commit

Wait for GPU free memory reporting to converge · 354ad925

Daniel Hiltgen authored May 09, 2024

The GPU drivers take a while to update their free memory reporting, so we need
to wait until the values converge with what we're expecting before proceeding
to start another runner in order to get an accurate picture.

354ad925

06 May, 2024 2 commits
- Skip scheduling cancelled requests, always reload unloaded runners (#4189) · c9f98622
  Jeffrey Morgan authored May 06, 2024
  
  c9f98622
- unload in critical section (#4187) · dfa2f32c
  Jeffrey Morgan authored May 05, 2024
  
  dfa2f32c
05 May, 2024 3 commits
- Centralize server config handling · f56aa200
  Daniel Hiltgen authored May 04, 2024
```
This moves all the env var reading into one central module
and logs the loaded config once at startup which should
help in troubleshooting user server logs
```
  f56aa200
- allocate a large enough kv cache for all parallel requests (#4162) · 942c9792
  Jeffrey Morgan authored May 05, 2024
  
  942c9792
- Make maximum pending request configurable · 20f6c065
  Daniel Hiltgen authored May 03, 2024
```
This also bumps up the default to be 50 queued requests
instead of 10.
```
  20f6c065
01 May, 2024 2 commits

log when the waiting for the process to stop to help debug when other tasks... · 63c76368

Mark Ward authored Apr 29, 2024

log when the waiting for the process to stop to help debug when other tasks execute during this wait.
expire timer clear the timer reference because it will not be reused.
close will clean up expireTimer if calling code has not already done this.

63c76368

fix runner expire during active use. Clearing the expire timer as it is used.... · f4a73d57

Mark Ward authored Apr 28, 2024

fix runner expire during active use.  Clearing the expire timer as it is used.  Allowing the finish to assign an expire timer so that the runner will expire after no use.

f4a73d57

28 Apr, 2024 1 commit

Fix concurrency for CPU mode · d6e3b645

Daniel Hiltgen authored Apr 28, 2024

Prior refactoring passes accidentally removed the logic to bypass VRAM
checks for CPU loads. This adds that back, along with test coverage.

This also fixes loaded map access in the unit test to be behind the mutex which was
likely the cause of various flakes in the tests.

d6e3b645

25 Apr, 2024 2 commits
- Reload model if `num_gpu` changes (#3920) · 00b0699c
  Jeffrey Morgan authored Apr 25, 2024
```
* reload model if `num_gpu` changes

* dont reload on -1

* fix tests
```
  00b0699c
- Adjust context size for parallelism · b123be5b
  Daniel Hiltgen authored Apr 25, 2024
  
  b123be5b
24 Apr, 2024 2 commits
- Restructure loading conditional chain · 36a6dacc
  Bryce Reitano authored Apr 24, 2024
  
  36a6dacc
- Move ggml loading to when we attempt fitting · 284e02be
  Bryce Reitano authored Apr 24, 2024
  
  284e02be
23 Apr, 2024 1 commit

Request and model concurrency · 34b9db5a

Daniel Hiltgen authored Mar 30, 2024

This change adds support for multiple concurrent requests, as well as
loading multiple models by spawning multiple runners. The default
settings are currently set at 1 concurrent request per model and only 1
loaded model at a time, but these can be adjusted by setting
OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS.

34b9db5a