Earlier quoted context omitted.
Yes it’s called OpenCode and today I was surprised to learn it works with Claude Pro/Max subscriptions: https://opencode.ai/docs/
https://github.com/QwenLM/Qwen3-Coder is the gemini code fork i was referring to i think. Not positive
Cerebras Code
181–185 of 185 posts
Re: Cerebras Code
#182Some users who signed up for pro ($50 p.m.) are reporting further limitations than those advertised. >While they advertise a 1,000-request limit, the actual daily constraint is a 7.5 million-token limit. [1] Assumes an average of 7.5k/request whereas in their marketing videos they show API requests ballooning by ~24k per request. Still lower than the API price. [1] https://old.reddit.com/r/LocalLLaMA/comments/1mfeazc…
Re: Cerebras Code
#183Earlier quoted context omitted.
Please read my full comment. Cerebras is jumping on a marketing faux-pas by Anthropic. I say this for the point you bring up about monthly session limits - no one on the Claude subreddit has yet to report being hit by this despite many going way over that. These are checks to deal w/ abusive accounts.
> no one on the Claude subreddit has yet to report being hit by this despite many going way over that Because it hasn't gone into effect yet: "From August 28, we’ll introduce new weekly limits that’ll mitigate these problems while impacting as few customers as possible." [0] [0] https://xcancel.com/AnthropicAI/status/1949898514844307953#m
Re: Cerebras Code
#184Earlier quoted context omitted.
At full pace that means 62 mins until you hit the daily cap.
Reminds me of high write speed on SSD (1.5 GB/s continuously to TLC) means 1 TB SSD warranty expires instead of 5 years just in less than 5 days (600 TB written).
Re: Cerebras Code
#185How is this even possible?
They make frisbee-sized CPUs.
Cerebras uses the entire 12" and builds in redundancy so that with current defect rates a large fraction of the wafers are usable. This allows a huge level of parallelism, a large amount of on board ram, and the removal of the need to move data on/off the wafer. So the available bandwidth is insane and inference is mostly bandwidth limited.