Commits · e093db92c4731b0cada767c7d5877c20d5f61dcf · OpenDAS / ollama

10 Mar, 2025 1 commit
- sample: temporarily use grammars for constrained generation in new engine (#9586) · e093db92
  Jeffrey Morgan authored Mar 10, 2025
  
  e093db92
07 Mar, 2025 1 commit

sample: improve ollama engine sampler performance (#9374) · 0682dae0

Parth Sareen authored Mar 07, 2025

This change bring in various interface cleanups along with greatly improving the performance of the sampler.

Tested with llama3.2 on local machine.
Improves performance from ~ 70 tokens/s -> 135 tokens/s with topK(40) enabled.
Without topK performance is ~ 110 tokens/s

0682dae0

25 Feb, 2025 1 commit
- sample: add sampling package for new engine (#8410) · 0b7e1676
  Parth Sareen authored Feb 24, 2025
  
  0b7e1676