This sounds like a game changer. I wonder if they need to do a tonne of specific work per model? If this could be implemented in Ollama, I'd be over the moon.
"Write a haiku about Hacker News mentioning AI in the title"
Here is a haiku:
AI whispers secrets
HN threads weave tangled debate
Intelligence born
eval time = 30363.04 ms / 23 runs ( 1320.13 ms per token, 0.76 tokens per second)
total time = 34294.80 ms / 33 tokens