1. 02 Aug, 2024 1 commit
  2. 14 Jun, 2024 2 commits
    • Daniel Hiltgen's avatar
      refined test timing · 68dfc623
      Daniel Hiltgen authored
      adjust timing on some tests so they don't timeout on small/slow GPUs
      68dfc623
    • Daniel Hiltgen's avatar
      Improve multi-gpu handling at the limit · 6fd04ca9
      Daniel Hiltgen authored
      Still not complete, needs some refinement to our prediction to understand the
      discrete GPUs available space so we can see how many layers fit in each one
      since we can't split one layer across multiple GPUs we can't treat free space
      as one logical block
      6fd04ca9
  3. 10 May, 2024 1 commit
  4. 06 May, 2024 1 commit
  5. 23 Apr, 2024 1 commit
    • Daniel Hiltgen's avatar
      Request and model concurrency · 34b9db5a
      Daniel Hiltgen authored
      This change adds support for multiple concurrent requests, as well as
      loading multiple models by spawning multiple runners. The default
      settings are currently set at 1 concurrent request per model and only 1
      loaded model at a time, but these can be adjusted by setting
      OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS.
      34b9db5a
  6. 01 Apr, 2024 1 commit
  7. 26 Mar, 2024 1 commit
  8. 25 Mar, 2024 1 commit
  9. 23 Mar, 2024 1 commit