Live data from Hacker News

Show HN: sllm – Split a GPU node with other developers, unlimited tokens

sllm.cloud

101–110 of 117 posts

Re: Show HN: sllm – Split a GPU node with other developers, unlimited tokens

#101
Thanks to everyone who shared feedback. We’re implementing it now.

Here’s what’s changed:

- We’ve removed the other LLMs for now and are focusing entirely on Qwen 3.5. We’ll bring back additional smaller models later, but most usage was already concentrated on Qwen 3.5.

- Pricing is now around $50. You get roughly 2× the throughput (61 tok/s vs. 31 tok/s, verified in testing), and it’s still unlimited. For context, that’s about 158M tokens per month. Comparable providers like Novita charge around $3.2 per million tokens, so this comes out to roughly 10% of typical token costs.

- Context size is now capped at 32K tokens. For the vast majority of use cases, this is more than sufficient.

Re: Show HN: sllm – Split a GPU node with other developers, unlimited tokens

#102

Pretty cool idea, but whats the stack behind this? As 15-25 tok/s seems a bit low as expected SoA for most providers is around 60 tok/s and quality of life dramatically improves above that.

15-25 was a rate based on oversubscription. Now it's 60 like others :).

Re: Show HN: sllm – Split a GPU node with other developers, unlimited tokens

#103

what is the main moat of your idea? privacy? otherwise it looks like a less flexible API compared to what chutes.ai or openrouter.ai providing. and they have TEE instances, which are more private. also why did u decide on launching V3 instead of some much more exciting models revealed recently like MiMo-V2-Pro or Arcee's Trinity Large?

You're right that we're less flexible than OpenRouter or Chutes. We don't let you hop between models per-request. If you want that, use those. If you want predictable cost and guaranteed throughput on one model, that's us.

On TEE: yeah, it's stronger, but it also adds cost and latency. We run dedicated hardware with no prompt logging and an isolated proxy. For most people who just don't want their data in someone's training set, that's enough. If your threat model is more serious than that, we're not the right choice.

On models: we are focusing on Qwen for now. We add based on demand. Would you actually use MiMo-V2-Pro or Trinity if we had them?

Re: Show HN: sllm – Split a GPU node with other developers, unlimited tokens

#106
I received an email mentioning that earlier cohorts are canceled.

Apparently, the earlier pricing was better for us as customers because I had the option to opt for a lower price, i.e $10 per month with a one month commitment and see how the platform evolves and then sign up for other models post testing as needed.

I am not sure how long this new cohort will take to fill now. A slightly better option (looking back) would have been to take multiple options from the customer list and start with the one that meets the threshold.

Re: Show HN: sllm – Split a GPU node with other developers, unlimited tokens

#107

[flagged]

The audience here is developers buying API access. They want to see the model, the price, and the throughput, not a hero image and three paragraphs about our mission. Marketing copy between a developer and that information is friction.

Re: Show HN: sllm – Split a GPU node with other developers, unlimited tokens

#109

[flagged]

We collect emails to notify you when the cohort fills or any important information such as cancellation. No one's selling your email.

Also, please read https://news.ycombinator.com/newsguidelines.html. HN is a community for thoughtful discussion.

Re: Show HN: sllm – Split a GPU node with other developers, unlimited tokens

#110

I received an email mentioning that earlier cohorts are canceled. Apparently, the earlier pricing was better for us as customers because I had the option to opt for a lower price, i.e $10 per month with a one month commitment and see how the platform evolves and then sign up for other models post testing as needed. I am not sure how long this new cohort will take to fill now. A slightly better option (looking back) w…

First, thanks for signing up early. It means a lot.

The $10/mo price needed 465 people to fill a cohort before we could turn on a single GPU. People signed up and churned while waiting, so we looked at the reservation pattern and determined 80 slots was optimal. This reflects in the new price and throughput.

We're considering a 1-week option so people can test it out before committing to a full month. Would that help?

Post reply on HN