Viewing profile — airgapstopgap
airgapstopgap
HN member- Joined
- Mon, Nov 28, 2022, 3:27 AM UTC
- HN karma
- 412
- Public activity
- 116 items
- HN profile
- View on Hacker News ↗
About airgapstopgap
No profile information was provided.
Recent public activity
-
comment
Comment #40071086
Can you point to anyone other than yourself who calls Indonesians black? Because I think otherwise it's not worthwhile discussing categorization and measurement in good faith with …
-
comment
Comment #39474843
Since you're here: have you considered moving to other, better generalist base models in the future? Particularly Deepseek or Mixtrals. Natural language foundation is important for…
-
comment
Comment #39087199
Note that we have no reason to believe that the underlying LLM inference process has suffered any setbacks. Obviously it has generated some logits. But the question is how is OpenA…
-
comment
Comment #38716423
This is not so surprising if you consider the fact that finetuning is extremely sparse and barely imparts any new knowledge to the model. The paper "Language Models are Super Mario…
-
comment
Comment #38622864
Intel aims to.
-
comment
Comment #38604436
The original paper by Shazeer suffices. What you are saying is in theory possible to do and may have been done in practice here, but in the general case MoE is trained from scratch…
-
comment
Comment #38602096
Mistral-small explicitly has inference costs of a 12.9b, but more than that, it's probably ran with batch size of 32 or higher. They'll worry more about offsetting training costs t…
-
comment
Comment #38601665
> It's not even close to a 45B model. They trained 8 different fine-tunes on the same base model. This means the 8 models differ only by a couple of layers and share the rest of th…
-
comment
Comment #38575816
> today we have no architecture or training methodology which would allow it to be possible. We clearly see that Mistral-7B is in some important, representative respects (eg coding…
-
comment
Comment #38443558
I do not even think any of this has much of impact on AGI timelines. Human brain cells are not a superior substrate for computing "intelligence". They just are what they are; indiv…
-
comment
Comment #38441046
Comments like this are incredibly grating. You condescend to the interlocutor for making a mistake which only exists in your own mistaken world model. Your confidence that neurons …
-
comment
Comment #38381882
…ETH Zurich is an illustrious research university that often cooperates with Deepmind and other hyped groups, they're right there at the frontier too, and have been for a very long…
-
comment
Comment #38357157
> murderous tendencies lurking beneath the surface …Where is that "beneath the surface"? Do you imagine a transformer has "thoughts" not dedicated to producing outputs? What is wit…
-
comment
Comment #38329155
> there is a possibility that for things like AI, with extra time comes the ability to better understand and build those defenses before they're needed. Or not, and damaging wrongh…
-
comment
Comment #37699603
Long-context tasks are not really the true gap between LLaMA and GPT series, but important result.
-
comment
Comment #37675876
Being authors of LLaMA is sufficient to argue they know how to train LLaMAs.
-
comment
Comment #37655755
Interested about your logic, what did you like about pre-LLM AGI? The "maximize utility function at any cost" feature? The single-minded focus on beating people in games? It's quit…
-
comment
Comment #37631755
Provable safety (not to confuse with security as in normal discussion of vulnerabilities) for general intelligence is a pipe dream because, putting things simply, undesirable reaso…
-
comment
Comment #37621391
Tegmark's thinking here is extremely shallow, discards the costs (opportunity costs and risks of stable dystopia) associated with this grandiose global project of dubious feasibili…
-
comment
Comment #37579567
I wonder if you have enough self-awareness to notice why your behavior here might be considered bizarre. No, people who point out that your government routinely and brazenly backdo…
-
comment
Comment #37250958
Llama-1-33B was trained on 40% more tokens than LLama-1-13B; this explained some of the disparity. This time around they both have the same data scale (2T pretraining + 500B code f…
-
comment
Comment #37097213
> Linux and Mac > Coming soon ... Ah well. Hopefully it is soon. Also, on behalf of all Apple Silicon Mac users, would be nice if the author looked into implementing Metal FlashAtt…
-
comment
Comment #37096638
This is an incredible achievement but there are strong reasons to suspect that stellarators are not and will never be plausible candidates for energy generation. For some more expe…
-
comment
Comment #37072139
Do you not consider that Huawei "executive's" detention (actual makes for a similar case against Canada? It was a purely political move, Meng Wanzhou was detained on grounds of a b…
-
comment
Comment #36901293
You are frustrated and this makes you act in a deliberately obtuse manner. There is a world of difference between "anyone who has worked with the guy" and "has worked with the guy …