Viewing profile — phazonoverload
phazonoverload
HN member- Joined
- Mon, Apr 26, 2021, 3:15 PM UTC
- HN karma
- 26
- Public activity
- 11 items
- HN profile
- View on Hacker News ↗
About phazonoverload
No profile information was provided.
Recent public activity
-
comment
Comment #49536010
I will give it a look!
-
comment
Comment #49535995
Coming back to this a few hours later, I've decided to add a section to explain this to the blog post. Thank you for flagging it.
-
comment
Comment #49535907
I haven't, but that's just because it really isn't my personal usage pattern.
-
comment
Comment #49535866
A literal manual typo. Good catch, will fix.
-
comment
Comment #49535841
The beauty of this is that you can just swap out the platform and everything remains as it's the same backend. You make a really good point, one that I haven't really considered, b…
-
comment
Comment #49535808
Hehe I really am just working it out as I go along - I promise it is fairly painless. Hugging Face allow you to specify your machine and then browse models that fit. And then you c…
-
comment
Comment #49534180
It is not
-
comment
Comment #49534170
Running it depends on RAM, which is what I wrote, bandwidth is important for speed. I chose my words carefully, but you are absolutely right.
-
comment
Comment #49534161
My perf sucks compared to yours. Added it to the post - same model averages 325 tok/s in processing prompts, and 34 tok/s in token generation. What am I doing wrong..?
-
comment
Comment #49534150
I'm the author - hello! Added to the post! Qwen averages 325 tok/s in processing prompts, and 34 tok/s in token generation. That isn't instant, but it's quick enough that I never r…
-
comment
Comment #49534148
I'm the author - hello! I talk about it in the blog post - knowing what's being run, knowing where it's being run, and not having anyone else control it.