Viewing profile — adroniser
adroniser
HN member- Joined
- Tue, Mar 21, 2023, 1:29 AM UTC
- HN karma
- 76
- Public activity
- 82 items
- HN profile
- View on Hacker News ↗
About adroniser
No profile information was provided.
Recent public activity
-
comment
Comment #47820050
what are you even yapping about.
-
comment
Comment #46667965
Adding the position vector is basic sure, but it's naive to think the model doesn't develop its own positional system bootstrapping on top of the barebones one.
-
comment
Comment #46321325
[flagged]
-
comment
Comment #45398079
fmri's are correlational nonsense (see Brainwashed, for example) and so are any "model introspection" tools.
-
comment
Comment #45252524
Of the papers submitted to a conference, it might be that reviewers don't offer suggestions that would significantly improve the quality of the work. Indeed the quality of reviews …
-
comment
Comment #45248298
So you think that this blog post would make it into any of the mainstream conferences? I doubt it.
-
comment
Comment #45247960
peer review would encourage less hand wavy language and more precise claims. They would penalize the authors for bringing up bizarre analogies to physics concepts for seemingly no …
- comment
-
comment
Comment #45050779
didn't hackers used to be for piracy?
-
comment
Comment #45050698
This suggests people should pre-register benchmarks. Because currently it feels like there is little incentive to publish benchmarks that models saturate.
-
comment
Comment #41526512
But there are lots of models available now that render much faster which are better quality than sora
-
comment
Comment #41241008
I completely agree this shit is so depressing. When I saw the AlphaProof paper I basically spent 3 days in mourning basically, because their approach was so simple.
-
comment
Comment #41240965
I think the whole paper is a satire lol.
-
comment
Comment #41240924
Does it really? If you want an LLM to edit code you need to feed it every single line of code in a prompt. Is it really that surprising that having just learnt it has been timed ou…
-
comment
Comment #41215958
I agree with you that transformers are probably not the architecture of choice. Not sure what that has to do with the viability of RL though.
-
comment
Comment #41215729
Hmm well the reason a pre-trained transformer is a fancy sentence completion engine is because that is what it is trained on, cross entropy loss on next token prediction. As I say,…
-
comment
Comment #41215107
But RL algorithms do implement things like curiosity to drive exploration?? https://arxiv.org/pdf/1810.12894 . Thinking to arbitrary depth sounds like Monte Carlo tree search? Whic…
-
comment
Comment #41208885
Isn't RL the algorithm we want basically?
-
comment
Comment #41208861
The distinction is that LLMs are not used for what they are trained for in this case. In the vast majority of cases someone using an LLM is not interested in what some mixture of o…
-
comment
Comment #41208833
How about you want to solve sudoku say.And you simply specify that you want the output to have unique numbers in each row, unique numbers in each column, and no unique number in an…
-
comment
Comment #41078017
"AlphaProof is a system that trains itself to prove mathematical statements in the formal language Lean. It couples a pre-trained language model with the AlphaZero reinforcement le…
-
comment
Comment #40907301
I did read your entire comment, and that is what prompted my response, because from my perspective your entire premise was based on LLMs failing at simple examples, and yet despite…
-
comment
Comment #40905285
If you're going to suggest something you think an LLM can't do I think at the very least as a show of good faith you should try it out. I've lost count of the number of times peopl…
-
comment
Comment #40811350
Yes but you are not taking an uncountable union. You are taking a finite union.
-
comment
Comment #40801058
To be clear, the construction given here violates the finite additivity property of measure. It's got nothing to do with the countable/uncountable additivity property.