Earlier quoted context omitted.
A) is just a sign, that further engineering, both model and framework, is required. Also, 50-60 words per second does not sound too bad. It amounts to 4M words per day. I imagine most people only type a thousand words per day. So with a single GPU you can handle 4000 users (assuming batch-style processing). B) This specific example actually shows ML toolchains are orders of magnitude better, than existing ones. Imagi…
On A: Using a p3.2xl for that task leaves you with a minimum ec2 cost of goods sold of 50 cents per user per month for the target task. To make a reasonable SaaS business out of this you would need to be charging a minimum of $5.00 per user month for the service that this GPU is providing, assuming all other tasks have a negligible impact on Cost of Goods Sold, and that you are able to spread the load optimally throu…
https://twitter.com/br_/status/979442438254166016
> "selling AWS at a loss" is crisp shorthand for a lot of startups' business models!