Commits · 7ca71a6b0fd4842e13cf9166f97ac2b58d4f874f · OpenDAS / ollama

10 May, 2024 1 commit

Daniel Hiltgen authored May 10, 2024

Under stress scenarios we're seeing OOMs so this should help stabilize
the allocations under heavy concurrency stress.

30a7d709

09 May, 2024 2 commits

Wait for GPU free memory reporting to converge · 354ad925

Daniel Hiltgen authored May 09, 2024

The GPU drivers take a while to update their free memory reporting, so we need
to wait until the values converge with what we're expecting before proceeding
to start another runner in order to get an accurate picture.

354ad925

Record more GPU information · 8727a9c1

Daniel Hiltgen authored May 07, 2024

This cleans up the logging for GPU discovery a bit, and can
serve as a foundation to report GPU information in a future UX.

8727a9c1

07 May, 2024 1 commit
- llm: add minimum based on layer size · 4736391b
  Michael Yang authored May 06, 2024
  
  4736391b
06 May, 2024 1 commit

Use our libraries first · 380378cc

Daniel Hiltgen authored May 05, 2024

Trying to live off the land for cuda libraries was not the right strategy.  We need to use the version we compiled against to ensure things work properly

380378cc

05 May, 2024 1 commit

Centralize server config handling · f56aa200

Daniel Hiltgen authored May 04, 2024

This moves all the env var reading into one central module
and logs the loaded config once at startup which should
help in troubleshooting user server logs

f56aa200

03 May, 2024 1 commit
- Skip PhysX cudart library · b1ad3a43
  Daniel Hiltgen authored May 03, 2024
```
For some reason this library gives incorrect GPU information, so skip it
```
  b1ad3a43
01 May, 2024 3 commits
- Support Fedoras standard ROCm location · e592e8fc
  Daniel Hiltgen authored May 01, 2024
  
  e592e8fc
- gpu: add 512MiB to darwin minimum, metal doesn't have partial offloading overhead (#4068) · f0c454ab
  Jeffrey Morgan authored May 01, 2024
  
  f0c454ab
- Add CUDA Driver API for GPU discovery · 089daaea
  Daniel Hiltgen authored Apr 30, 2024
```
We're seeing some corner cases with cudart which might be resolved by
switching to the driver API which comes bundled with the driver package
```
  089daaea
29 Apr, 2024 1 commit
- Fix relative path lookup · 7b59d177
  Daniel Hiltgen authored Apr 29, 2024
  
  7b59d177
26 Apr, 2024 1 commit
- also look at cwd as a root for windows runners (#3959) · aad8d128
  Jeffrey Morgan authored Apr 26, 2024
  
  aad8d128
24 Apr, 2024 1 commit
- AMD gfx patch rev is hex · 0d6687f8
  Daniel Hiltgen authored Apr 24, 2024
```
Correctly handle gfx90a discovery
```
  0d6687f8
23 Apr, 2024 2 commits

Move nested payloads to installer and zip file on windows · 058f6cd2

Daniel Hiltgen authored Apr 23, 2024

Now that the llm runner is an executable and not just a dll, more users are facing
problems with security policy configurations on windows that prevent users
writing to directories and then executing binaries from the same location.
This change removes payloads from the main executable on windows and shifts them
over to be packaged in the installer and discovered based on the executables location.
This also adds a new zip file for people who want to "roll their own" installation model.

058f6cd2

Request and model concurrency · 34b9db5a

Daniel Hiltgen authored Mar 30, 2024

This change adds support for multiple concurrent requests, as well as
loading multiple models by spawning multiple runners. The default
settings are currently set at 1 concurrent request per model and only 1
loaded model at a time, but these can be adjusted by setting
OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS.

34b9db5a

16 Apr, 2024 2 commits
- scale graph based on gpu count · 26df6747
  Michael Yang authored Apr 16, 2024
  
  26df6747
- darwin: no partial offloading if required memory greater than system · 41a272de
  Michael Yang authored Apr 16, 2024
  
  41a272de
10 Apr, 2024 1 commit
- partial offloading · 7e33a017
  Michael Yang authored Apr 05, 2024
  
  7e33a017
01 Apr, 2024 6 commits

Refined min memory from testing · 1f11b525
Daniel Hiltgen authored Apr 01, 2024

1f11b525

Release gpu discovery library after use · 526d4eb2

Daniel Hiltgen authored Mar 30, 2024

Leaving the cudart library loaded kept ~30m of memory
pinned in the GPU in the main process.  This change ensures
we don't hold GPU resources when idle.

526d4eb2

Safeguard for noexec · 0a74cb31

Daniel Hiltgen authored Mar 28, 2024

We may have users that run into problems with our current
payload model, so this gives us an escape valve.

0a74cb31

Detect too-old cuda driver · 10ed1b62
Daniel Hiltgen authored Mar 28, 2024
```
"cudart init failure: 35" isn't particularly helpful in the logs.
```
10ed1b62

Switch back to subprocessing for llama.cpp · 58d95cc9

Daniel Hiltgen authored Mar 14, 2024

This should resolve a number of memory leak and stability defects by allowing
us to isolate llama.cpp in a separate process and shutdown when idle, and
gracefully restart if it has problems. This also serves as a first step to be
able to run multiple copies to support multiple models concurrently.

58d95cc9

update memory calcualtions · 91b3e4d2
Michael Yang authored Mar 18, 2024
```
count each layer independently when deciding gpu offloading
```
91b3e4d2

28 Mar, 2024 1 commit
- Update troubleshooting link · f31f2bed
  Michael Yang authored Mar 28, 2024
  
  f31f2bed
25 Mar, 2024 1 commit
- add support for libcudart.so for CUDA devices (adds Jetson support) · dfc6721b
  Jeremy authored Mar 25, 2024
  
  dfc6721b
20 Mar, 2024 1 commit

Better tmpdir cleanup · 74788b48

Daniel Hiltgen authored Mar 13, 2024

If expanding the runners fails, don't leave a corrupt/incomplete payloads dir
We now write a pid file out to the tmpdir, which allows us to scan for stale tmpdirs
and remove this as long as there isn't still a process running.

74788b48

12 Mar, 2024 2 commits

Fix iGPU detection for linux · 82b0c7c2

Daniel Hiltgen authored Mar 12, 2024

This fixes a few bugs in the new sysfs discovery logic. iGPUs are now
correctly identified by their <1G VRAM reported. the sysfs IDs are off
by one compared to what HIP wants due to the CPU being reported
in amdgpu, but HIP only cares about GPUs.

82b0c7c2

fix gpu_info_cuda.c compile warning (#3077) · 51578d85
mofanke authored Mar 13, 2024

51578d85

11 Mar, 2024 1 commit

Avoid rocm runner and dependency clash · bc13da2b

Daniel Hiltgen authored Mar 11, 2024

Putting the rocm symlink next to the runners is risky.  This moves
the payloads into a subdir to avoid potential clashes.

bc13da2b

10 Mar, 2024 1 commit

Add ollama executable peer dir for rocm · 00ec2693

Daniel Hiltgen authored Mar 10, 2024

This allows people who package up ollama on their own to place
the rocm dependencies in a peer directory to the ollama executable
much like our windows install flow.

00ec2693

09 Mar, 2024 2 commits

tidy cleanup logs · 0bd0f4a2
Jeffrey Morgan authored Mar 09, 2024

0bd0f4a2

Finish unwinding idempotent payload logic · 4a5c9b80

Daniel Hiltgen authored Mar 08, 2024

The recent ROCm change partially removed idempotent
payloads, but the ggml-metal.metal file for mac was still
idempotent.  This finishes switching to always extract
the payloads, and now that idempotentcy is gone, the
version directory is no longer useful.

4a5c9b80

07 Mar, 2024 2 commits

Revamp ROCm support · 6c5ccb11

Daniel Hiltgen authored Feb 15, 2024

This refines where we extract the LLM libraries to by adding a new
OLLAMA_HOME env var, that defaults to `~/.ollama` The logic was already
idempotenent, so this should speed up startups after the first time a
new release is deployed. It also cleans up after itself.

We now build only a single ROCm version (latest major) on both windows
and linux. Given the large size of ROCms tensor files, we split the
dependency out. It's bundled into the installer on windows, and a
separate download on windows. The linux install script is now smart and
detects the presence of AMD GPUs and looks to see if rocm v6 is already
present, and if not, then downloads our dependency tar file.

For Linux discovery, we now use sysfs and check each GPU against what
ROCm supports so we can degrade to CPU gracefully instead of having
llama.cpp+rocm assert/crash on us. For Windows, we now use go's windows
dynamic library loading logic to access the amdhip64.dll APIs to query
the GPU information.

6c5ccb11

Allow setting max vram for workarounds · be330174

Daniel Hiltgen authored Mar 06, 2024

Until we get all the memory calculations correct, this can provide
and escape valve for users to workaround out of memory crashes.

be330174

29 Feb, 2024 1 commit
- fix: print usedMemory size right (#2827) · fa2f2b35
  tylinux authored Mar 01, 2024
  
  fa2f2b35
25 Feb, 2024 1 commit

Determine max VRAM on macOS using `recommendedMaxWorkingSetSize` (#2354) · a189810d

peanut256 authored Feb 26, 2024

* read iogpu.wired_limit_mb on macOS

Fix for https://github.com/ollama/ollama/issues/1826

* improved determination of available vram on macOS

read the recommended maximal vram on macOS via Metal API

* Removed macOS-specific logging

* Remove logging from gpu_darwin.go

* release Core Foundation object

fixes a possible memory leak

a189810d

17 Feb, 2024 1 commit
- Harden AMD driver lookup logic · 9754c6d9
  Daniel Hiltgen authored Feb 16, 2024
```
It looks like the version file doesnt exist on older(?) drivers
```
  9754c6d9
12 Feb, 2024 1 commit

Detect AMD GPU info via sysfs and block old cards · 6d84f075

Daniel Hiltgen authored Feb 11, 2024

This wires up some new logic to start using sysfs to discover AMD GPU
information and detects old cards we can't yet support so we can fallback to CPU mode.

6d84f075

28 Jan, 2024 1 commit
- Don't disable GPUs on arm without AVX · 15562e88
  Daniel Hiltgen authored Jan 28, 2024
```
AVX is an x86 feature, so ARM should be excluded from
the check.
```
  15562e88