Viewing profile — scottcha
scottcha
HN member- Joined
- Tue, Jan 18, 2011, 6:47 PM UTC
- HN karma
- 964
- Public activity
- 155 items
- HN profile
- View on Hacker News ↗
About scottcha
Recent public activity
-
comment
Comment #49022095
We use vllm as it generally has the best ecosystem support. Parameters are largely dependent on what type of requests you are serving (concurrency, input/output ratios, cached hit …
-
comment
Comment #49012078
I run an AI platform and we need to tokenize fast and early to make a lot of decisions on the subsequent steps (things like routing, rate limiting and such). Its really important t…
-
comment
Comment #48900940
Hi, I'm co-founder of Neuralwatt. While there aren't other providers selling by energy we also produce the same tokens stats you get at others (input, output, cached) and you can c…
-
comment
Comment #48880842
They do have a growing amount of Scope 1 emissions (emissions from their on site sources) which originally was primarily on site diesel but due to grid interconnect delays have bee…
-
comment
Comment #48720674
Thanks for the feedback! Our primary focus is charging by energy, for token pricing we really just try to be close to the market. That being said I'll take a look at our token pric…
-
comment
Comment #48713760
Hi I'm the CTO of neuralwatt, would love to hear your feedback on what your experience was. Feel free to email me scott@neuralwatt.com. Also for GLM5.2 we run the FP8 quantization …
-
comment
Comment #48510694
Yes way better. We host both and while qwen3.6 is over 100tps we usually can do glm around that too.
-
comment
Comment #48503142
I use glm5.1 plus pi with a few customized skills and am very happy with it. I hadn’t touched my Claude 5x plan for a couple of weeks but opened it back up in Claude code when fabl…
-
comment
Comment #48183569
I use claude code and pi.dev side by side most days and i'm mostly choosing pi for most work in last couple of weeks.
-
comment
Comment #47642315
Pretty cool idea, but whats the stack behind this? As 15-25 tok/s seems a bit low as expected SoA for most providers is around 60 tok/s and quality of life dramatically improves ab…
-
comment
Comment #47634300
I think there was a clarification posted on Reddit that said Claude Agents SDK didn't apply for now.
-
comment
Comment #47632093
I use OpenCode and have just started using Nanoclaw with ClaudeCode (my coworker has a post coming on this) and sometimes ClaudeCode with Claude Code Router. I do a range of small …
-
comment
Comment #47630293
Mine are pretty unique since we optimize the energy for and run an inference service api so forces me to dogfood alot of different options.
-
comment
Comment #47630183
Yes GLM5 and KimiK2.5 are pretty close replacements for sonnet.
-
comment
Comment #47616204
I switch between Claude Code (Opus/Sonnet) and Qwen (OpenCode, OpenClaw) multiple times throughout the day and Qwen 3.5 is really nice. I do also use KimiK2.5 and GLM5 pretty often…
-
comment
Comment #47587532
We offer multiple SOA models at https://portal.neuralwatt.com at very generous pricing since we have options to bill per kWh instead of per token. Recipes for your favorite tools h…
-
comment
Comment #47313724
I actually built this analysis while I worked at Microsoft so I 100% agree. Doing the work at the platform level is the way to go and you can actually make a significant impact wit…
-
comment
Comment #47172452
There have been a few questions about the state of Show HN lately. Was actually interested in this post but I see all the OPs responses to questions are Dead? I do see its a new ac…
- story
-
comment
Comment #46699091
That is a pretty good article although the one factor not mentioned that we see that has a huge impact on energy is batch size but that would be hard to estimate with the data he h…
-
comment
Comment #45096461
Neuralwatt | https://neuralwatt.com | REMOTE (US – Seattle/Denver/Boulder metros only) | Full-time | $180k–$220k DOE Energy is the #1 constraint in new datacenter buildouts. Neural…
-
comment
Comment #44767967
Neuralwatt | https://neuralwatt.com | REMOTE (US – Seattle/Denver/Boulder metros only) | Full-time | $180k–$220k DOE Energy is the #1 constraint in new datacenter buildouts. Neural…
-
comment
Comment #44658434
I’ve asked that question on linked in to the Cerebras team a couple times and haven’t ever received a response. There is system max tdp values posted online but I’m not sure you ca…
-
comment
Comment #44253063
Turns out there is multiple publications associating tachycardia and other heart symptoms with long covid. https://pmc.ncbi.nlm.nih.gov/articles/PMC8356730/
-
comment
Comment #44250028
Maybe I’m a statistical anomaly or maybe I just don’t know the baseline occurrence rate for this stuff but I have 3 close acquaintances two of which are this persons age or younger…