Commits · ceb0e26e5e8d60228eaa4e04d85869cb19d823c3 · orangecat / ollama

"docs/source/experiment/webportal.rst" did not exist on "abd164c2598d4cf19a081b4e5c1070de7bea8386"

24 Apr, 2024 1 commit
- Move ggml loading to when we attempt fitting · 284e02be
  Bryce Reitano authored Apr 24, 2024
  
  284e02be
23 Apr, 2024 1 commit

Request and model concurrency · 34b9db5a

Daniel Hiltgen authored Mar 30, 2024

This change adds support for multiple concurrent requests, as well as
loading multiple models by spawning multiple runners. The default
settings are currently set at 1 concurrent request per model and only 1
loaded model at a time, but these can be adjusted by setting
OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS.

34b9db5a