Commits · e9f7f3602961d2b0beaff27144ec89301c2173ca · OpenDAS / ollama

13 Jul, 2024 1 commit
- llm: looser checks for minimum memory (#5677) · ef98803d
  Jeffrey Morgan authored Jul 13, 2024
  
  ef98803d
12 Jul, 2024 1 commit
- fix: quant err message (#5616) · 10e76882
  Josh authored Jul 11, 2024
  
  10e76882
11 Jul, 2024 3 commits

llm: avoid loading model if system memory is too small (#5637) · c4cf8ad5

Jeffrey Morgan authored Jul 11, 2024



* llm: avoid loading model if system memory is too small

* update log

* Instrument swap free space

On linux and windows, expose how much swap space is available
so we can take that into consideration when scheduling models

* use `systemSwapFreeMemory` in check

---------
Co-authored-by: Daniel Hiltgen <daniel@ollama.com>

c4cf8ad5

sched: only error when over-allocating system memory (#5626) · 791650dd
Jeffrey Morgan authored Jul 11, 2024

791650dd
llm: dont link cuda with compat libs (#5621) · efbf41ed
Jeffrey Morgan authored Jul 10, 2024

efbf41ed

10 Jul, 2024 4 commits
- chatglm graph · 5a739ff4
  Michael Yang authored Jul 10, 2024
  
  5a739ff4
- remove `GGML_CUDA_FORCE_MMQ=on` from build (#5588) · 4e262eb2
  Jeffrey Morgan authored Jul 10, 2024
  
  4e262eb2
- Bump ROCm on windows to 6.1.2 · 1f50356e
  Daniel Hiltgen authored Jul 10, 2024
```
This also adjusts our algorithm to favor our bundled ROCm.
I've confirmed VRAM reporting still doesn't work properly so we
can't yet enable concurrency by default.
```
  1f50356e
- Remove duplicate merge glitch · 22c81f62
  Daniel Hiltgen authored Jul 10, 2024
  
  22c81f62
09 Jul, 2024 1 commit

Statically link c++ and thread lib · b51e3b63

Daniel Hiltgen authored Jul 09, 2024

This makes sure we statically link the c++ and thread library on windows
to avoid unnecessary runtime dependencies on non-standard DLLs

b51e3b63

08 Jul, 2024 1 commit
- Workaround broken ROCm p2p copy · 0bacb300
  Daniel Hiltgen authored Jul 05, 2024
```
Enable the build flag for llama.cpp to use CPU copy for multi-GPU scenarios.
```
  0bacb300
07 Jul, 2024 4 commits
- llm: remove ambiguous comment when putting upper limit on predictions to avoid... · 53da2c69
  Jeffrey Morgan authored Jul 07, 2024
```
llm: remove ambiguous comment when putting upper limit on predictions to avoid infinite generation (#5535)
```
  53da2c69
- llm: allow gemma 2 to context shift (#5534) · d8def1ff
  Jeffrey Morgan authored Jul 07, 2024
  
  d8def1ff
- Update llama.cpp submodule to `a8db2a9c` (#5530) · 571dc619
  Jeffrey Morgan authored Jul 07, 2024
  
  571dc619
- llm: print caching notices in debug only (#5533) · 0e09c380
  Jeffrey Morgan authored Jul 07, 2024
  
  0e09c380
06 Jul, 2024 8 commits
- llm: add `-DBUILD_SHARED_LIBS=off` to common cpu cmake flags (#5520) · 4607c706
  Jeffrey Morgan authored Jul 06, 2024
  
  4607c706
- release: remove unwanted mingw dll.a files · a08f20d9
  jmorganca authored Jul 06, 2024
  
  a08f20d9
- Revert "llm: only statically link libstdc++" · 6cea0360
  jmorganca authored Jul 06, 2024
```
This reverts commit 5796bfc4.
```
  6cea0360
- llm: only statically link libstdc++ · 5796bfc4
  jmorganca authored Jul 06, 2024
  
  5796bfc4
- llm: statically link pthread and stdc++ dependencies in windows build · f1a379aa
  jmorganca authored Jul 06, 2024
  
  f1a379aa
- llm: add `GGML_STATIC` flag to windows static lib · 9ae14699
  jmorganca authored Jul 06, 2024
  
  9ae14699
- llm: add `COMMON_DARWIN_DEFS` to arm static build (#5513) · e0348d3f
  Jeffrey Morgan authored Jul 05, 2024
  
  e0348d3f
- llm: fix missing dylibs by restoring old build behavior on Linux and macOS (#5511) · 2cc854f8
  Jeffrey Morgan authored Jul 05, 2024
```
* Revert "fix cmake build (#5505)"

This reverts commit 4fd5f352.

* llm: fix missing dylibs by restoring old build behavior

* crlf -> lf
```
  2cc854f8
05 Jul, 2024 7 commits
- llm: put back old include dir (#5507) · 5304b765
  Jeffrey Morgan authored Jul 05, 2024
```
* llm: put back old include dir

* llm: update link paths for old submodule commits
```
  5304b765
- fix cmake build (#5505) · 4fd5f352
  Jeffrey Morgan authored Jul 05, 2024
  
  4fd5f352
- fix model reloading · ac7a842e
  Michael Yang authored Jul 03, 2024
```
ensure runtime model changes (template, system prompt, messages,
options) are captured on model updates without needing to reload the
server
```
  ac7a842e
- fix typo in cgo directives in `llm.go` (#5501) · 78fb33dd
  Jeffrey Morgan authored Jul 05, 2024
  
  78fb33dd
- update llama.cpp submodule to `d7fd29f` (#5475) · 8f8e736b
  Jeffrey Morgan authored Jul 05, 2024
  
  8f8e736b
- Use slot with cached prompt instead of least recently used (#5492) · d89454de
  Jeffrey Morgan authored Jul 05, 2024
```
* Use common prefix to select slot

* actually report `longest`
```
  d89454de
- Fix assert on small embedding inputs (#5491) · e9188e97
  Jeffrey Morgan authored Jul 05, 2024
```
* Fix assert on small embedding inputs

* Update llm/patches/09-pooling.diff
```
  e9188e97
04 Jul, 2024 1 commit
- fix error detection by limiting model loading error parsing (#5472) · 4d71c559
  Jeffrey Morgan authored Jul 03, 2024
  
  4d71c559
03 Jul, 2024 3 commits

Return Correct Prompt Eval Count Regardless of Cache Prompt (#5371) · 3b5a4a77

royjhan authored Jul 03, 2024

* openai compatibility

* Revert "openai compatibility"

This reverts commit d3f98a811e00fc497d889c8c45b0cfec5b64690c.

* remove erroneous subtraction of prompt cache

3b5a4a77

Fix corner cases on tmp cleaner on mac · 0e982bc1

Daniel Hiltgen authored Jul 03, 2024

When ollama is running a long time, tmp cleaners can remove the
runners. This tightens up a few corner cases on arm macs where
we failed with "server cpu not listed in available servers map[]"

0e982bc1

Fix clip model loading with unicode paths · 6298f498

Daniel Hiltgen authored Jul 03, 2024

On windows, if the model dir contained unicode characters
clip models would fail to load.  This fixes the file name
handling in clip.cpp to support utf16 on windows.

6298f498

01 Jul, 2024 2 commits
- error · 33a65e3b
  Josh Yan authored Jul 01, 2024
  
  33a65e3b
- Switch use_mmap to a pointer type · 97c9e117
  Daniel Hiltgen authored Jun 28, 2024
```
This uses nil as undefined for a cleaner implementation.
```
  97c9e117
29 Jun, 2024 1 commit
- Do not shift context for sliding window models (#5368) · 717f7229
  Jeffrey Morgan authored Jun 28, 2024
```
* Do not shift context for sliding window models

* truncate prompt > 2/3 tokens

* only target gemma2
```
  717f7229
27 Jun, 2024 2 commits
- gemma2 graph · de2163da
  Michael Yang authored Jun 27, 2024
  
  de2163da
- llm: architecture patch (#5316) · 4d311eb7
  Jeffrey Morgan authored Jun 26, 2024
  
  4d311eb7
25 Jun, 2024 1 commit

llm: speed up gguf decoding by a lot (#5246) · cb42e607

Blake Mizerany authored Jun 24, 2024

Previously, some costly things were causing the loading of GGUF files
and their metadata and tensor information to be VERY slow:

  * Too many allocations when decoding strings
  * Hitting disk for each read of each key and value, resulting in a
    not-okay amount of syscalls/disk I/O.

The show API is now down to 33ms from 800ms+ for llama3 on a macbook pro
m3.

This commit also prevents collecting large arrays of values when
decoding GGUFs (if desired). When such keys are encountered, their
values are null, and are encoded as such in JSON.

Also, this fixes a broken test that was not encoding valid GGUF.

cb42e607