Viewing profile — germanjoey
germanjoey
HN member- Joined
- Tue, Mar 27, 2018, 9:27 PM UTC
- HN karma
- 225
- Public activity
- 48 items
- HN profile
- View on Hacker News ↗
About germanjoey
No profile information was provided.
Recent public activity
-
comment
Comment #48416331
IMO "bugs per commit" is even worse than that, because, in addition to what you say, it also hides the extraordinary spike of commit activity of a project that had previously been …
-
comment
Comment #43989756
TBH, the 2x-4x improvement over a naive implementation that they're bragging about sounded kinda pathetic to me! I mean, it depends greatly on the kernel itself and the target arch…
-
comment
Comment #42704562
This is really incredible, thank you!
-
comment
Comment #42407405
Sambanova's RDU is a dataflow processor being used for ML/AI workloads! It's amazing and actually works.
-
comment
Comment #42179744
Pretty amazing speed, especially considering this is bf16. But how many racks is this using? The used 4 racks for 70B, so this, what, at least 24? A whole data center for one model…
-
comment
Comment #42003207
the title says "Cerebras Trains Llama Models"...
-
comment
Comment #41942732
They said in the announcement that they've implemented speculative decoding, so that might have a lot to do with it. A big question is what they're using as their draft model; ther…
-
comment
Comment #41745848
Simply increasing processing power for the AI isn't enough. Gameplay mechanics are intimately related to the capabilities of the AI. For example, when they redesigned combat around…
-
comment
Comment #41631662
How are you verifying accuracy for your JAX port of Llama 3.1? IMHO, the main reason to use pytorch is actually that the original model used pytorch. What can seem to be identical …
-
comment
Comment #41385859
Looks like some kind of power play... Originally discussed here: https://news.ycombinator.com/item?id=41234180
- story
-
comment
Comment #40360334
Is there a demo of a model visualized using this somewhere? Even if it's just a short video... it's hard to tell what it's like from screenshots.
-
comment
Comment #39965094
cost effective in what sense? groq doesn't achieve high efficiency, only low latency. but that's not done in a cost-effective way. compare sambanova achieving the same performance …
-
comment
Comment #39866205
We're showing off our 1.05T param Composition of Experts LLM! It's 150 experts running on 1 node consisting of 8 SN40L RDU chips. Each of our nodes has a huge amount of DDR attache…
- story
-
comment
Comment #37638630
Sambanova just launched something similar to what you're describing. It's a demo of their new chip running a 1T param MoE model 150 7B llama2s, each retrained to be an expert in a …
- story
-
comment
Comment #35157016
welp, This report focuses on the capabilities, limitations, and safety properties of GPT-4. GPT-4 is a Transformer-style model [33 ] pre-trained to predict the next token in a docu…
-
comment
Comment #35156664
How big is this model? (i.e., how many parameters?) I can't find this anywhere.
-
comment
Comment #34807244
I worked with the author for a couple of years, pre- and post- acquisition, and I have to admit that he drove me somewhat crazy sometimes too. Leaving that aside, I also had an imm…
-
comment
Comment #34254510
What's the new performance process?
-
comment
Comment #33485006
> You don't introduce more coupling, you don't the coupling that already exists. This is true at the code level. But at the system-design level, this documentation is the extra cou…
-
comment
Comment #33085492
It is interesting reading that second paragraph many years later. Most of the things that Steve Yegge brags about that Google "does right" (e.g. how they do recruiting, their engin…
-
comment
Comment #31950291
Great post; this is how I felt about it too.
-
comment
Comment #30805898
The article (or, rather, the commentary in the link above on the article) talks about the fallacious notion of "market cap" in regards to cryptocurrencies. That is to say, e.g., mu…