Commits · 755b4e4fc291366595ed7bfb37c2a91ff5834df8 · OpenDAS / ollama

19 Jun, 2024 1 commit
- Revert "gpu: add env var for detecting Intel oneapi gpus (#5076)" · 755b4e4f
  Wang,Zhe authored Jun 19, 2024
```
This reverts commit 163cd3e7.
```
  755b4e4f
17 Jun, 2024 2 commits
- Fix a build warning (#5096) · 4ad0d4d6
  Lei Jitang authored Jun 18, 2024
```
Signed-off-by: Lei Jitang <leijitang@outlook.com>
```
  4ad0d4d6
- gpu: add env var for detecting Intel oneapi gpus (#5076) · 163cd3e7
  Jeffrey Morgan authored Jun 16, 2024
```
* gpu: add env var for detecting intel oneapi gpus

* fix build error
```
  163cd3e7
16 Jun, 2024 1 commit
- Add some more debugging logs for intel discovery · fd1e6e05
  Daniel Hiltgen authored Jun 16, 2024
```
Also removes an unused overall count variable
```
  fd1e6e05
15 Jun, 2024 1 commit
- gpu: Fix build warning · 225f0d12
  Lei Jitang authored Jun 15, 2024
```
Signed-off-by: Lei Jitang <leijitang@outlook.com>
```
  225f0d12
14 Jun, 2024 10 commits

Centralize GPU configuration vars · 6be309e1

Daniel Hiltgen authored May 08, 2024

This should aid in troubleshooting by capturing and reporting the GPU
settings at startup in the logs along with all the other server settings.

6be309e1

Workaround gfx900 SDMA bugs · da3bf233

Daniel Hiltgen authored May 31, 2024

Implement support for GPU env var workarounds, and leverage
this for the Vega RX 56 which needs
HSA_ENABLE_SDMA=0 set to work properly

da3bf233

review comments and coverage · 6f351bf5
Daniel Hiltgen authored Jun 05, 2024

6f351bf5
Refine CPU load behavior with system memory visibility · fc37c192
Daniel Hiltgen authored Jun 03, 2024

fc37c192

Reintroduce nvidia nvml library for windows · 434dfe30

Daniel Hiltgen authored Jun 03, 2024

This library will give us the most reliable free VRAM reporting on windows
to enable concurrent model scheduling.

434dfe30

Refactor intel gpu discovery · 4e2b7e18
Daniel Hiltgen authored May 29, 2024

4e2b7e18

Improve multi-gpu handling at the limit · 6fd04ca9

Daniel Hiltgen authored May 18, 2024

Still not complete, needs some refinement to our prediction to understand the
discrete GPUs available space so we can see how many layers fit in each one
since we can't split one layer across multiple GPUs we can't treat free space
as one logical block

6fd04ca9

Refine GPU discovery to bootstrap once · 43ed358f

Daniel Hiltgen authored May 15, 2024

Now that we call the GPU discovery routines many times to
update memory, this splits initial discovery from free memory
updating.

43ed358f

Use DRM driver for VRAM info for amd · b32ebb4f

Daniel Hiltgen authored May 14, 2024

The amdgpu drivers free VRAM reporting omits some other apps, so leverage the
upstream DRM driver which keeps better tabs on things

b32ebb4f

Revert "Limit GPU lib search for now (#4777)" · efac4886
Daniel Hiltgen authored Jun 03, 2024
```
This reverts commit 476fb8e8.
```
efac4886

13 Jun, 2024 1 commit
- Actually skip PhysX on windows · aac36763
  Daniel Hiltgen authored Jun 13, 2024
  
  aac36763
04 Jun, 2024 3 commits
- lint windows · e919f681
  Michael Yang authored May 22, 2024
  
  e919f681
- lint linux · bf7edb0d
  Michael Yang authored May 22, 2024
  
  bf7edb0d
- lint · e40145a3
  Michael Yang authored May 21, 2024
  
  e40145a3
02 Jun, 2024 1 commit
- Limit GPU lib search for now (#4777) · 476fb8e8
  Jeffrey Morgan authored Jun 01, 2024
```
* fix oneapi errors on windows 10
```
  476fb8e8
24 May, 2024 2 commits
- Move envconfig and consolidate env vars (#4608) · 4cc3be30
  Patrick Devine authored May 24, 2024
  
  4cc3be30
- support ollama run on Intel GPUs · fd5971be
  Wang,Zhe authored May 24, 2024
  
  fd5971be
10 May, 2024 1 commit

Bump VRAM buffer back up · 30a7d709

Daniel Hiltgen authored May 10, 2024

Under stress scenarios we're seeing OOMs so this should help stabilize
the allocations under heavy concurrency stress.

30a7d709

09 May, 2024 2 commits

Wait for GPU free memory reporting to converge · 354ad925

Daniel Hiltgen authored May 09, 2024

The GPU drivers take a while to update their free memory reporting, so we need
to wait until the values converge with what we're expecting before proceeding
to start another runner in order to get an accurate picture.

354ad925

Record more GPU information · 8727a9c1

Daniel Hiltgen authored May 07, 2024

This cleans up the logging for GPU discovery a bit, and can
serve as a foundation to report GPU information in a future UX.

8727a9c1

07 May, 2024 1 commit
- llm: add minimum based on layer size · 4736391b
  Michael Yang authored May 06, 2024
  
  4736391b
06 May, 2024 1 commit

Use our libraries first · 380378cc

Daniel Hiltgen authored May 05, 2024

Trying to live off the land for cuda libraries was not the right strategy.  We need to use the version we compiled against to ensure things work properly

380378cc

05 May, 2024 1 commit

Centralize server config handling · f56aa200

Daniel Hiltgen authored May 04, 2024

This moves all the env var reading into one central module
and logs the loaded config once at startup which should
help in troubleshooting user server logs

f56aa200

03 May, 2024 1 commit
- Skip PhysX cudart library · b1ad3a43
  Daniel Hiltgen authored May 03, 2024
```
For some reason this library gives incorrect GPU information, so skip it
```
  b1ad3a43
01 May, 2024 3 commits
- Support Fedoras standard ROCm location · e592e8fc
  Daniel Hiltgen authored May 01, 2024
  
  e592e8fc
- gpu: add 512MiB to darwin minimum, metal doesn't have partial offloading overhead (#4068) · f0c454ab
  Jeffrey Morgan authored May 01, 2024
  
  f0c454ab
- Add CUDA Driver API for GPU discovery · 089daaea
  Daniel Hiltgen authored Apr 30, 2024
```
We're seeing some corner cases with cudart which might be resolved by
switching to the driver API which comes bundled with the driver package
```
  089daaea
29 Apr, 2024 1 commit
- Fix relative path lookup · 7b59d177
  Daniel Hiltgen authored Apr 29, 2024
  
  7b59d177
26 Apr, 2024 1 commit
- also look at cwd as a root for windows runners (#3959) · aad8d128
  Jeffrey Morgan authored Apr 26, 2024
  
  aad8d128
24 Apr, 2024 1 commit
- AMD gfx patch rev is hex · 0d6687f8
  Daniel Hiltgen authored Apr 24, 2024
```
Correctly handle gfx90a discovery
```
  0d6687f8
23 Apr, 2024 2 commits

Move nested payloads to installer and zip file on windows · 058f6cd2

Daniel Hiltgen authored Apr 23, 2024

Now that the llm runner is an executable and not just a dll, more users are facing
problems with security policy configurations on windows that prevent users
writing to directories and then executing binaries from the same location.
This change removes payloads from the main executable on windows and shifts them
over to be packaged in the installer and discovered based on the executables location.
This also adds a new zip file for people who want to "roll their own" installation model.

058f6cd2

Request and model concurrency · 34b9db5a

Daniel Hiltgen authored Mar 30, 2024

This change adds support for multiple concurrent requests, as well as
loading multiple models by spawning multiple runners. The default
settings are currently set at 1 concurrent request per model and only 1
loaded model at a time, but these can be adjusted by setting
OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS.

34b9db5a

16 Apr, 2024 2 commits
- scale graph based on gpu count · 26df6747
  Michael Yang authored Apr 16, 2024
  
  26df6747
- darwin: no partial offloading if required memory greater than system · 41a272de
  Michael Yang authored Apr 16, 2024
  
  41a272de
10 Apr, 2024 1 commit
- partial offloading · 7e33a017
  Michael Yang authored Apr 05, 2024
  
  7e33a017