- 11 Nov, 2024 1 commit
-
-
Evan authored
-
- 10 Nov, 2024 1 commit
-
-
Arhan Busam authored
-
- 08 Nov, 2024 3 commits
-
-
Jesse Gross authored
If we get a request with a zero length image, it will result in an out-of-bounds error when we pass the data to the image encoder.
-
Edward J. Schwartz authored
-
Daniel Hiltgen authored
-
- 07 Nov, 2024 5 commits
-
-
Daniel Hiltgen authored
This should have been in #7347 but was overlooked.
-
Daniel Hiltgen authored
This enables the workaround code only for windows which should help windows users with muliple AMD GPUs
-
Daniel Hiltgen authored
Some users are reporting crashes during nvcuda.dll initialization on windows. This should help narrow down where things are going bad.
-
Daniel Hiltgen authored
Bring consistency with the old generate script behavior
-
Daniel Hiltgen authored
On linux nvcc isn't automatically linking to the same cuda version.
-
- 06 Nov, 2024 3 commits
-
-
Jesse Gross authored
-
Jesse Gross authored
Now that server.cpp is gone, we don't need to keep passing arguments that were only ignored and only kept for compatibility.
-
Jesse Gross authored
The Go runner does not have a problem with supporting parallel requests for most multimodal models. Now that we won't be potentially falling back to server.cpp, this restriction can be lifted. However, the new mllama model can't support parallel requests, so we will need to keep a restriction for that.
-
- 05 Nov, 2024 4 commits
-
-
RAPID ARCHITECT authored
added reddit rate below hexabot, ollama powered reddit search and analysis with streamlit for the intervace
-
Daniel Hiltgen authored
One potential failure mode is an empty file which bubbles up as an EOF error, leading to all pulls and listing operations failing. Instead, continue and warn about the corrupt manifest. This also allows re-pulling the corrupt manifest to repair the system.
-
Jesse Gross authored
Currently we assume that images take 768 tokens of context size for the purposes of clipping old messages that exceed the context window. However, our mllama implementation stores the full image embedding in a single token. As a result, there is significant waste of context space. Ideally, we would handle this more generically and have the implementation report the number of tokens. However, at the moment this would just result in a similar set of 'if' conditions in the runner plus APIs to report it back. So for now, we just keep this simple.
-
Med Marrouchi authored
-
- 04 Nov, 2024 6 commits
-
-
Daniel Hiltgen authored
Avoid excessive log spew and make consistent with chat logging
-
Daniel Hiltgen authored
-
Daniel Hiltgen authored
Github actions matrix strategy can't access env settings
-
Michael Yang authored
update llama3.2 vision memory estimation
-
Daniel Hiltgen authored
-
suncloudsmoon authored
-
- 02 Nov, 2024 4 commits
-
-
Daniel Hiltgen authored
The runtime and management libraries may not always have identical ordering, so use the device UUID to correlate instead of ID.
-
Daniel Hiltgen authored
This leverages caching, and some reduced installer scope to try to speed up builds. It also tidies up some windows build logic that was only relevant for the older generate/cmake builds.
-
Jesse Gross authored
Check for NULL return values from llama.cpp in more places and convert them into Go errors, which should make debugging easier in the future rather than having hidden surprises in our data structures.
-
Jesse Gross authored
Mllama has large embeddings (100 MB per image) and each embedding is represented as 1 token when passed to llama.cpp. Batches are pre- allocated for the size of the tokens times the batch size, so this results in allocations of over 50 GB at the default batch size. On some systems, these mallocs will fail. Since an image is represented as a single token and mllama doesn't support more than 1 image per request, we only need to allocate a batch size of 1, which is much more reasonable. In addition, for non-multimodal models, we don't need to allocate the embedding batches at all. Fixes #7464
-
- 01 Nov, 2024 3 commits
-
-
Michael Yang authored
-
Michael Yang authored
-
Daniel Hiltgen authored
-
- 31 Oct, 2024 2 commits
-
-
Jesse Gross authored
Currently if an input has embeddings at any point then we will set cross attention to true from the beginning. This means that any tokens before the embeddings are sent will incorrectly have cross attention layers applied. This only sets cross attention when we have an embedding, either previously in this sequence or in the cache. It also makes cross attention capable of supporting parallelism at the runner level, though the mllama implementation doesn't support that yet.
-
Daniel Hiltgen authored
* Give unicode test more time to run Some slower GPUs (or partial CPU/GPU loads) can take more than the default 30s to complete this test * Give more time for concurrency test CPU inference can be very slow under stress
-
- 30 Oct, 2024 6 commits
-
-
Daniel Hiltgen authored
Until we have full NUMA support, this adjusts the default thread selection algorithm to count up the number of performance cores across all sockets.
-
Jesse Gross authored
-Update mllama to take the cross attention state as embeddings in a batch, more similar to how Llava handles it. This improves integration with the input cache. -Pass locations in a prompt for embeddings using tags similar to Llava. -Abstract interface to vision models so the main runner accesses Clip and Mllama similarly Co-authored-by:Michael Yang <mxyng@pm.me>
-
Daniel Hiltgen authored
This will no longer error if built with regular gcc on windows. To help triage issues that may come in related to different compilers, the runner now reports the compier used by cgo.
-
Daniel Hiltgen authored
* Remove llama.cpp submodule and shift new build to top * CI: install msys and clang gcc on win Needed for deepseek to work properly on windows
-
Daniel Hiltgen authored
-
Daniel Hiltgen authored
* windows: Support alt install paths Advanced users are leveraging innosetup's /DIR switch to target an alternate location, but we get confused by things not existing in the LocalAppData dir. This also hardens the server path lookup code for a future attempt to unify with a ./bin prefix * Fit and finish improvements for windows app Document alternate install location instructions for binaries and model. Pop up progress UI for upgrades (automatic, with cancel button). Expose non-default port in menu to disambiguate mutiple instances. Set minimum Windows version to 10 22H2
-
- 29 Oct, 2024 2 commits
-
-
Patrick Devine authored
-
Daniel Hiltgen authored
* Switch over to clang for deepseek on windows The patch for deepseek requires clang on windows. gcc on windows has a buggy c++ library and can't handle the unicode characters * Fail fast with wrong compiler on windows Avoid users mistakenly building with GCC when we need clang
-