Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of time reading Read about 5M tokens - Output is awesome, super fast as you expect from the 1500t/sec I think that's correct - Tool call is failing more than say DS4, which leads to time wasted on retries (complex tools like browser control for example) - Shell comm…
> "Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy" I don't believe Cerebras has a cached input pricing? They don't list one on the model page: https://inference-docs.cerebras.ai/models/qwen-3.8-27b edit: See the sibling discussion, https://news.ycombinator.com/item?id=49554520#49555094 ( "Input tokens, whether served from the cache or processed fresh,…
I wonder if they will do that with sol ultrafast!