Commits · b80661e8c78e115ed9b41391c87fdb7f1a7f69ec · OpenDAS / ollama

"vscode:/vscode.git/clone" did not exist on "5a1d2722ac47a192089b9b820ab7fd831411866a"

11 Mar, 2024 1 commit
- relay load model errors to the client (#3065) · b80661e8
  Bruce MacDonald authored Mar 11, 2024
  
  b80661e8
07 Mar, 2024 1 commit

Daniel Hiltgen authored Feb 15, 2024

This refines where we extract the LLM libraries to by adding a new
OLLAMA_HOME env var, that defaults to `~/.ollama` The logic was already
idempotenent, so this should speed up startups after the first time a
new release is deployed. It also cleans up after itself.

We now build only a single ROCm version (latest major) on both windows
and linux. Given the large size of ROCms tensor files, we split the
dependency out. It's bundled into the installer on windows, and a
separate download on windows. The linux install script is now smart and
detects the presence of AMD GPUs and looks to see if rocm v6 is already
present, and if not, then downloads our dependency tar file.

For Linux discovery, we now use sysfs and check each GPU against what
ROCm supports so we can degrade to CPU gracefully instead of having
llama.cpp+rocm assert/crash on us. For Windows, we now use go's windows
dynamic library loading logic to access the amdhip64.dll APIs to query
the GPU information.

6c5ccb11

20 Feb, 2024 1 commit
- update llama.cpp submodule to `66c1968f7` (#2618) · 4613a080
  Jeffrey Morgan authored Feb 20, 2024
  
  4613a080
09 Feb, 2024 1 commit

Shutdown faster · 66807615

Daniel Hiltgen authored Feb 08, 2024

Make sure that when a shutdown signal comes, we shutdown quickly instead
of waiting for a potentially long exchange to wrap up.

66807615

01 Feb, 2024 1 commit
- use `llm.ImageData` · f11bf074
  Jeffrey Morgan authored Jan 31, 2024
  
  f11bf074
29 Jan, 2024 1 commit
- remove unknown `CPPFLAGS` option · 2e06ed01
  Jeffrey Morgan authored Jan 28, 2024
  
  2e06ed01
25 Jan, 2024 1 commit
- Fix clearing kv cache between requests with the same prompt (#2186) · a64570dc
  Jeffrey Morgan authored Jan 25, 2024
```
* Fix clearing kv cache between requests with the same prompt

* fix powershell script
```
  a64570dc
22 Jan, 2024 2 commits
- Refine debug logging for llm · 730dcfcc
  Daniel Hiltgen authored Jan 22, 2024
```
This wires up logging in llama.cpp to always go to stderr, and also
turns up logging if OLLAMA_DEBUG is set.
```
  730dcfcc
- Debug logging on init failure · 27a2d5af
  Daniel Hiltgen authored Jan 22, 2024
  
  27a2d5af
21 Jan, 2024 1 commit
- Unlock mutex when failing to load model (#2117) · 89c4aee2
  Jeffrey Morgan authored Jan 20, 2024
  
  89c4aee2
18 Jan, 2024 1 commit
- Mechanical switch from log to slog · fedd705a
  Daniel Hiltgen authored Jan 18, 2024
```
A few obvious levels were adjusted, but generally everything mapped to "info" level.
```
  fedd705a
17 Jan, 2024 1 commit
- Add multiple CPU variants for Intel Mac · 1b249748
  Daniel Hiltgen authored Jan 12, 2024
```
This also refines the build process for the ext_server build.
```
  1b249748
16 Jan, 2024 1 commit
- do not cache prompt (#2018) · a897e833
  Bruce MacDonald authored Jan 16, 2024
```
- prompt cache causes inferance to hang after some time
```
  a897e833
13 Jan, 2024 1 commit
- Fix intel mac build · 2ecb2472
  Daniel Hiltgen authored Jan 13, 2024
```
Make sure we're building an x86 ext_server lib when cross-compiling
```
  2ecb2472
11 Jan, 2024 2 commits

Always dynamically load the llm server library · 39928a42

Daniel Hiltgen authored Jan 09, 2024

This switches darwin to dynamic loading, and refactors the code now that no
static linking of the library is used on any platform

39928a42

Build multiple CPU variants and pick the best · d88c527b

Daniel Hiltgen authored Jan 07, 2024

This reduces the built-in linux version to not use any vector extensions
which enables the resulting builds to run under Rosetta on MacOS in
Docker. Then at runtime it checks for the actual CPU vector
extensions and loads the best CPU library available

d88c527b

08 Jan, 2024 1 commit

Offload layers to GPU based on new model size estimates (#1850) · 08f1e189

Jeffrey Morgan authored Jan 08, 2024



* select layers based on estimated model memory usage

* always account for scratch vram

* dont load +1 layers

* better estmation for graph alloc

* Update gpu/gpu_darwin.go
Co-authored-by: Bruce MacDonald <brucewmacdonald@gmail.com>

* Update llm/llm.go
Co-authored-by: Bruce MacDonald <brucewmacdonald@gmail.com>

* Update llm/llm.go

* add overhead for cuda memory

* Update llm/llm.go
Co-authored-by: Bruce MacDonald <brucewmacdonald@gmail.com>

* fix build error on linux

* address comments

---------
Co-authored-by: Bruce MacDonald <brucewmacdonald@gmail.com>

08f1e189

07 Jan, 2024 1 commit
- dont use `-Wall` in static build (#1833) · 5feec959
  Jeffrey Morgan authored Jan 07, 2024
  
  5feec959
04 Jan, 2024 1 commit
- Code shuffle to clean up the llm dir · 77d96da9
  Daniel Hiltgen authored Jan 04, 2024
  
  77d96da9
03 Jan, 2024 1 commit
- fix: relay request opts to loaded llm prediction (#1761) · 0b3118e0
  Bruce MacDonald authored Jan 03, 2024
  
  0b3118e0
02 Jan, 2024 2 commits

Switch windows build to fully dynamic · d966b730

Daniel Hiltgen authored Dec 23, 2023

Refactor where we store build outputs, and support a fully dynamic loading
model on windows so the base executable has no special dependencies thus
doesn't require a special PATH.

d966b730

Refactor how we augment llama.cpp · 9a70aecc

Daniel Hiltgen authored Dec 22, 2023

This changes the model for llama.cpp inclusion so we're not applying a patch,
but instead have the C++ code directly in the ollama tree, which should make it
easier to refine and update over time.

9a70aecc

27 Dec, 2023 1 commit
- enable `cache_prompt` by default · d4ebdadb
  Jeffrey Morgan authored Dec 27, 2023
  
  d4ebdadb
22 Dec, 2023 2 commits
- Add Cache flag to api (#1642) · 10da41d6
  K0IN authored Dec 22, 2023
  
  10da41d6
- Fix CPU performance on hyperthreaded systems · 325d7498
  Daniel Hiltgen authored Dec 21, 2023
```
The default thread count logic was broken and resulted in 2x the number
of threads as it should on a hyperthreading CPU
resulting in thrashing and poor performance.
```
  325d7498
21 Dec, 2023 1 commit

Revive windows build · d9cd3d96

Daniel Hiltgen authored Dec 20, 2023

The windows native setup still needs some more work, but this gets it building
again and if you set the PATH properly, you can run the resulting exe on a cuda system.

d9cd3d96

20 Dec, 2023 1 commit

Revamp the dynamic library shim · 7555ea44

Daniel Hiltgen authored Dec 20, 2023

This switches the default llama.cpp to be CPU based, and builds the GPU variants
as dynamically loaded libraries which we can select at runtime.

This also bumps the ROCm library to version 6 given 5.7 builds don't work
on the latest ROCm library that just shipped.

7555ea44

19 Dec, 2023 5 commits
- Fix darwin intel build · 6558f94e
  Daniel Hiltgen authored Dec 19, 2023
  
  6558f94e
- Refine build to support CPU only · 1b991d0b
  Daniel Hiltgen authored Dec 13, 2023
```
If someone checks out the ollama repo and doesn't install the CUDA
library, this will ensure they can build a CPU only version
```
  1b991d0b
- Bump llama.cpp to b1662 and set n_parallel=1 · 9adca7f7
  Daniel Hiltgen authored Dec 14, 2023
  
  9adca7f7
- Adapted rocm support to cgo based llama.cpp · 35934b2e
  Daniel Hiltgen authored Nov 29, 2023
  
  35934b2e
- Add cgo implementation for llama.cpp · d4cd6957
  Daniel Hiltgen authored Nov 13, 2023
```
Run the server.cpp directly inside the Go runtime via cgo
while retaining the LLM Go abstractions.
```
  d4cd6957