1. 24 Sep, 2025 1 commit
  2. 19 Sep, 2025 1 commit
  3. 16 Sep, 2025 1 commit
  4. 05 Sep, 2025 2 commits
  5. 02 Sep, 2025 1 commit
  6. 22 Aug, 2025 1 commit
  7. 21 Aug, 2025 1 commit
  8. 20 Aug, 2025 1 commit
  9. 19 Aug, 2025 1 commit
  10. 18 Aug, 2025 1 commit
  11. 15 Aug, 2025 2 commits
  12. 14 Aug, 2025 1 commit
  13. 13 Aug, 2025 1 commit
  14. 11 Aug, 2025 1 commit
  15. 07 Aug, 2025 2 commits
  16. 01 Aug, 2025 1 commit
  17. 30 Jul, 2025 1 commit
  18. 28 Jul, 2025 1 commit
  19. 25 Jul, 2025 1 commit
  20. 22 Jul, 2025 1 commit
  21. 16 Jul, 2025 1 commit
  22. 08 Jul, 2025 1 commit
  23. 07 Jul, 2025 1 commit
  24. 03 Jul, 2025 1 commit
  25. 13 Jun, 2025 1 commit
  26. 29 May, 2025 1 commit
  27. 23 May, 2025 1 commit
  28. 19 May, 2025 1 commit
    • Graham King's avatar
      feat: Support multiple models on single ingress node (#1127) · aeb79e62
      Graham King authored
      We can now do this:
      
      - Node 1:
      
      ```
      dynamo-run in=http out=dyn
      ```
      
      - Node 2 and 3, two instances of component 'backend' in the nemotron_ultra pipeline:
      
      ```
      dynamo-run in=dyn://nemotron_ultra.backend.generate out=vllm /data/models/NemotronUltra
      ```
      
      - Node 4 and 5, two instances of the 'backend' component in nemotron_super pipeline:
      
      ```
      dynamo-run in=dyn://nemotron_super.backend.generate out=vllm /data/models/NemotronSuper
      ```
      
      The ingress node will discover all four instances and route correctly. We have been planning for this for a long time now.
      
      As part of this auto-discovery is now always `out=dyn`, with no extra URL parts. Previously it could only route to a single pipeline.
      
      Also:
      - Refactor endpoint / instance naming now that I understand them
      - Fix removing models when their instance stops.
      aeb79e62
  29. 16 May, 2025 1 commit
  30. 09 May, 2025 2 commits
  31. 29 Apr, 2025 1 commit
    • Graham King's avatar
      chore: Split PushRouter from Client (#817) · a1a10365
      Graham King authored
      In a distributed system we don't know if the remote workers need pre-processing done ingress-side or not. Previously Client required us to decide this before discovering the remote endpoints, which was fine because pre-processing was worker-side.
      
      As part of moving pre-processing back to ingress-side we need to split this into two steps:
      - Client discovers the endpoints, and (later PR) will fetch their Model Deployment Card.
      - PushRouter will use the Model Deployment Card to decide if they need pre-processing or not, which affects the types of the generic parameters.
      
      Part of #743
      a1a10365
  32. 25 Apr, 2025 2 commits
    • Harrison Saturley-Hall's avatar
    • Graham King's avatar
      chore: Publish Model Deployment Card to NATS (#799) · d346782c
      Graham King authored
      This will allow an ingress-side pre-processor to see it without needing a model checkout.
      
      Currently pre-processing is done in the worker, which has access to the model deployment card ("MDC") files (`config.json`, `tokenizer.json` and `tokenizer_config.json`) locally. We want to move the pre-processor to the ingress side to support KV routing. That requires ingress side (i.e the HTTP server), on a different machine than the worker to be able to see those three files.
      
      To support that this PR makes the worker upload the contents of those files to the NATS object store, and publishes the MDC with those NATS urls to the key-value store. 
      
      The key-value store has an interface so any store (nats, etcd, redis, etc) can be supported. Implementations for memory and NATS are provided.
      
      Fetching the MDC from the store, doing pre-processing ingress side, and publishing a card backed by a GGUF, are all for a later commit.
      
      Part of #743 
      d346782c
  33. 09 Apr, 2025 1 commit
  34. 04 Apr, 2025 1 commit
    • Graham King's avatar
      chore: Upgrade Rust to 1.86 (#518) · e99aa1e1
      Graham King authored
      Also upgrade the cargo resolver to v3, the default.
      
      New clippy lints:
      - `next_back()` instead of `last()` for a double-ended iterator. That avoids walking the whole list.
      - ` repeat_n` instead of `repeat.take`. That avoids cloning.
      - Doc indenting
      e99aa1e1
  35. 31 Mar, 2025 1 commit