Commits · f29b167e1af8a7c0e8a15044584826432aff76d8 · OpenDAS / ollama

04 Sep, 2024 1 commit

Use cuda v11 for driver 525 and older (#6620) · f29b167e

Daniel Hiltgen authored Sep 03, 2024

It looks like driver 525 (aka, cuda driver 12.0) has problems with the cuda v12 library
we compile against, so run v11 on those older drivers if detected.

f29b167e

19 Aug, 2024 2 commits
- Review comments · f9e31da9
  Daniel Hiltgen authored Aug 15, 2024
  
  f9e31da9
- Add cuda v12 variant and selection logic · 4fe3a556
  Daniel Hiltgen authored Jun 13, 2024
```
Based on compute capability and driver version, pick
v12 or v11 cuda variants.
```
  4fe3a556
04 Jun, 2024 1 commit
- lint linux · bf7edb0d
  Michael Yang authored May 22, 2024
  
  bf7edb0d
23 Apr, 2024 1 commit

Request and model concurrency · 34b9db5a

Daniel Hiltgen authored Mar 30, 2024

This change adds support for multiple concurrent requests, as well as
loading multiple models by spawning multiple runners. The default
settings are currently set at 1 concurrent request per model and only 1
loaded model at a time, but these can be adjusted by setting
OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS.

34b9db5a