Viewing profile — am17an
am17an
HN member- Joined
- Sun, Sep 04, 2016, 1:46 AM UTC
- HN karma
- 209
- Public activity
- 84 items
- HN profile
- View on Hacker News ↗
About am17an
Recent public activity
-
comment
Comment #49002618
It's expensive now, I expect once it is with inference providers it will be really dirt cheap. Then it would be truly be a "bicycle for the mind", which Fable promised to be except…
-
comment
Comment #48988694
Your username indeed checks out
-
comment
Comment #48988402
I think we're saying the same thing? I said they are explicitly trained on these tasks, not that they are some separate models during programming RL, or business tasks RL. > The co…
-
comment
Comment #48988180
Have you looked at the what data companies (e.g. Scale, Mercor) hire for? Why do you think Meta records their employees every keystroke/mousestroke/eye-movement? EDIT: just re-read…
-
comment
Comment #48988041
> nor are they trained for each task individually. They are explicitly trained for each task individually.
-
comment
Comment #48893709
Having looked at his code, I doubt this.
-
comment
Comment #48722200
llama 3? Are you from 2023?
-
comment
Comment #48289945
Who is their right minds would be wedded to an identity of saying "No"? Code quality puritans are annoying but if they do their job right they actually speed-up the development pro…
-
comment
Comment #48268384
Fair enough. I agree with you - although DS4 Pro is a GPT 5 class model which scores 46% on ARC-AGI-2[^1]. It's behind by maybe 9 months, I think it's still good enough for a lot o…
-
comment
Comment #48265485
Did you read the OP when he's exactly chiding the model you're glazing?
-
comment
Comment #48258679
This Claude front end skill is now soon to be slop.
-
comment
Comment #48231806
People doing economics with the cloud GPUs, of course cloud GPUs are going cheaper. But also, is generating tokens all you do with your computer? I can play games on DGX spark and …
-
comment
Comment #48092196
Local models embody the hacker spirit, constant Claude glazing is spiritually incompatible with tinkering. Don't upload your spirit to the cloud.
-
comment
Comment #47942748
No I mean more expensive, i.e. you're consuming vastly more tokens.
-
comment
Comment #47900256
The cost being reduced is the cost of your labour. Tokens are only getting more expensive.
-
comment
Comment #47898406
Sounds exhausting. Are your revenue numbers up?
-
comment
Comment #47748076
Don’t underestimate the march of technology. Just look at your phone, it has more FLOPS than there were in the entire world 40 years ago.
-
comment
Comment #47575793
Thank you, there are two things I would like to point out: 1) Google releasing something probably means they don't see it as important. 4-bit KV-cache quantization has been known f…
-
comment
Comment #47564911
There are techniques which already achieve great compression of the cache at 4 bit, eg using hadamard transforms. Going from 4 bit to 3 bit isn’t the great leap people expect this …
-
comment
Comment #47440213
Welp, back to pip
-
comment
Comment #47386216
Working in open source, I've now heard a wide variety of disabilities that people have and they have to be aided by an LLM for writing even descriptions of their PRs.
-
comment
Comment #47367387
You can still run larger MoE models using expert weight off-loading to the CPU for token generation. They are by and large useable, I get ~50 toks/second on a kimi linear 48B (3B a…
-
comment
Comment #47264696
Sure. “Tell me a joke”
-
comment
Comment #47209063
I was referring to the 35B version. It is surprisingly good for its size. You can use it for implementation tasks without it going off the rails
-
comment
Comment #47208777
Damn I’m jealous that they figured out how to pay their contributors. I’ve been toiling away for free