"lib/vscode:/vscode.git/clone" did not exist on "f2ba58e568a91489c6bdf024cef64cab1a9e9f77"
  1. 11 Jun, 2025 1 commit
  2. 09 Jun, 2025 3 commits
  3. 06 Jun, 2025 1 commit
  4. 04 Jun, 2025 4 commits
  5. 03 Jun, 2025 1 commit
  6. 02 Jun, 2025 2 commits
  7. 30 May, 2025 3 commits
  8. 29 May, 2025 8 commits
  9. 28 May, 2025 3 commits
  10. 27 May, 2025 1 commit
  11. 24 May, 2025 1 commit
  12. 23 May, 2025 4 commits
  13. 22 May, 2025 4 commits
    • Graham King's avatar
      feat(dynamo-run): Allow setting KV cache block size (#1175) · 183f2b32
      Graham King authored
      Example:
      ```
      dynamo-run out=<engine> <model> --kv-cache-block-size 64
      ```
      
      In a distributed system this goes on the worker node and is propagated to ingress via the model deployment card.
      
      Previously hard coded to 16, which is now the default.
      
      - Load context_length from model. Closes #1172
      - Store context length and KV cache block size in Model Deployment Card #1170
      183f2b32
    • Graham King's avatar
      fix: Fix race condition in kv_router unit test (#1174) · 3bde1e45
      Graham King authored
      Removed the hard coded sleeps, explained what we're testing.
      
      Closes https://github.com/ai-dynamo/dynamo/issues/1132
      
      The race condition is that `apply_event` sends a message on a channel, it does not directly apply the event. At some later point the tokio runtime schedules the task running the channel receiver, which applies the event. If that had not happened yet the test would fail.
      3bde1e45
    • jthomson04's avatar
      feat: Various KVBM improvements (#1134) · 5d5080ba
      jthomson04 authored
      5d5080ba
    • Graham King's avatar
      feat(dynamo-run): Allow setting context-length (#1157) · 6d5da821
      Graham King authored
      Llama 4 has a very large context length (aka n_ctx, model_max_length, max_model_len), and vllm won't start unless it can allocate enough KV cache for the entire context.
      
      Allow passing `--context-length <N>` to `dynamo-run` to limit it so long-context models will fit.
      
      Future todo:
      - Restrict every request's `max_tokens` to below the context length. Our pre-processor should do this by setting stop_conditions.max_tokens. mistralrs engine wrapper must do it itself because it does not use the pre-processor.
      - mistralrs and llamacpp currently have a hard-coded max context length if one is not provided on the command line. Change those to be the model's built-in max, read from the GGUF or tokenizer_config.json.
      6d5da821
  14. 21 May, 2025 3 commits
  15. 20 May, 2025 1 commit