Viewing profile — x_may
x_may
HN member- Joined
- Tue, Aug 27, 2024, 1:08 PM UTC
- HN karma
- 65
- Public activity
- 23 items
- HN profile
- View on Hacker News ↗
About x_may
No profile information was provided.
Recent public activity
-
comment
Comment #48278330
Yeah, but this is partly due to there being a shortage of entry level GPUs for consumers. NVIDIA has literally stopped manufacturing them. There are massive numbers of data centre …
-
comment
Comment #47516067
KV cache compression, so how much memory the model needs to use for extending its context. Does not affect the weight size.
-
comment
Comment #47122070
Isn’t there also basically 0 American DRAM?
-
comment
Comment #47035741
The 80/20 rule always wins
-
comment
Comment #47025663
I just wanted deterministic outputs and was curious how you were doing it. Sounds like probably temp = 0, which major providers no longer offer. Thanks for your response.
-
comment
Comment #47023497
Wait sorry how did you use and expose seeds? That’s the most interesting part of your post
-
comment
Comment #46852054
It might have been explicitly targeted, but they did say that there were older versions of Notepad ++ with ""insufficient update verification controls" so it might have just been t…
-
comment
Comment #46667916
I believe the what the parent comment was referring to is the advice not to praise character, but instead praise hard work. “You’re so smart” leaves room for failure when they enco…
-
comment
Comment #44482226
I think it’s also largely driven by the apparently cheapness of turning the CapEX of server buying to the OpEX of cloud renting. Less up front investment and auditing/access contro…
-
comment
Comment #44106380
Unfortunate name collision on that one
-
comment
Comment #44065932
Obviously its not at the scale of the top auto-regressive models yet but there are some OSS models https://github.com/dllm-reasoning/d1
-
comment
Comment #43727132
It may be that it was time for the hardware that was previously running Arxiv to be retired and this is just another Capex -> Opex decision being made by so many tech companies. I'…
-
comment
Comment #42566034
I believe they are using scalable TTC. The o3 announcement released accuracy numbers for high and low compute usage, which I feel would be hard to do in the same model without TTC.…
-
comment
Comment #42560826
The LMSYS leaderboards are crowdsourced and would be hard to fake, it showing a pretty strong performance in terms of human preference.
-
comment
Comment #42559392
Captcha solvers as a service are already well developed. The end result is going full circle to in person applications only.
-
comment
Comment #42559383
Tragedy of the commons at work once again
-
comment
Comment #42441650
There’s black sand! Volcanic sand from Iceland is perfectly black and would be a great way to distinguish them
-
comment
Comment #41887314
I think right now they lose more money with each user. But maybe their value lies in training data
-
comment
Comment #41847333
Not as much as meta, no. But AI21 labs is partnered with Amazon and did a ~$200M funding round last year IIRC so still plenty of funds for training big models
-
comment
Comment #41847312
As another commenter said, this has no GGUF because it’s partially mamba based which is unsupported in llama.cpp
-
comment
Comment #41689296
We’ve all had moments like that
-
comment
Comment #41616465
Check out Nougat from meta
-
comment
Comment #41367124
Sell them and invest the money