1. 20 Jun, 2024 3 commits
  2. 19 Jun, 2024 4 commits
  3. 17 Jun, 2024 3 commits
  4. 16 Jun, 2024 1 commit
  5. 15 Jun, 2024 1 commit
  6. 14 Jun, 2024 10 commits
  7. 13 Jun, 2024 1 commit
  8. 04 Jun, 2024 3 commits
  9. 02 Jun, 2024 1 commit
  10. 24 May, 2024 2 commits
  11. 10 May, 2024 1 commit
    • Daniel Hiltgen's avatar
      Bump VRAM buffer back up · 30a7d709
      Daniel Hiltgen authored
      Under stress scenarios we're seeing OOMs so this should help stabilize
      the allocations under heavy concurrency stress.
      30a7d709
  12. 09 May, 2024 2 commits
    • Daniel Hiltgen's avatar
      Wait for GPU free memory reporting to converge · 354ad925
      Daniel Hiltgen authored
      The GPU drivers take a while to update their free memory reporting, so we need
      to wait until the values converge with what we're expecting before proceeding
      to start another runner in order to get an accurate picture.
      354ad925
    • Daniel Hiltgen's avatar
      Record more GPU information · 8727a9c1
      Daniel Hiltgen authored
      This cleans up the logging for GPU discovery a bit, and can
      serve as a foundation to report GPU information in a future UX.
      8727a9c1
  13. 07 May, 2024 1 commit
  14. 06 May, 2024 1 commit
    • Daniel Hiltgen's avatar
      Use our libraries first · 380378cc
      Daniel Hiltgen authored
      Trying to live off the land for cuda libraries was not the right strategy.  We need to use the version we compiled against to ensure things work properly
      380378cc
  15. 05 May, 2024 1 commit
    • Daniel Hiltgen's avatar
      Centralize server config handling · f56aa200
      Daniel Hiltgen authored
      This moves all the env var reading into one central module
      and logs the loaded config once at startup which should
      help in troubleshooting user server logs
      f56aa200
  16. 03 May, 2024 1 commit
  17. 01 May, 2024 3 commits
  18. 29 Apr, 2024 1 commit