Viewing profile — mxwsn
mxwsn
HN member- Joined
- Fri, Sep 15, 2017, 3:05 PM UTC
- HN karma
- 1,662
- Public activity
- 237 items
- HN profile
- View on Hacker News ↗
About mxwsn
No profile information was provided.
Recent public activity
-
comment
Comment #48884181
This is the same reasoning behind why Yann Lecun thought test-time scaling would not work for LLMs: compounding error. Instead, the more tokens LLMs use, the better their performan…
- story
-
comment
Comment #48248956
No, there are more training tokens than parameters in LLMs. They are in the classical first descent setting.
-
comment
Comment #48072268
> Here’s a thought experiment: suppose that a mathematician solved a major problem by having a long exchange with an LLM in which the mathematician played a useful guiding role but…
-
comment
Comment #48056545
Great summary. The fact that the auto encoding task is not grounded in thoughts, and their initial training on guessed internal thoughts, raise serious concerns on faithfulness. Fe…
- comment
-
comment
Comment #48041064
Diffusion and flow matching models generate samples by iterative denoising. Iterative denoising means passing input to the neural network, running a forward pass, and taking the ou…
-
comment
Comment #47912574
How do you know that width scaling has been the driving force of improvement?
-
comment
Comment #46640794
The Jacobian is first derivatives, but for a function mapping N to M dimensions. It's the first derivative of every output wrt every input, so it will be an N x M matrix. The gradi…
-
comment
Comment #45511981
Wow! The title suggests introductory material, but in my opinion this has strong potential to win test of time awards for research.
-
comment
Comment #45434262
That's really interesting. What if they RAG search related videos from the prompt, and condition on that to generate? That might explain fidelity like this
-
comment
Comment #45342496
Why is not the diffusion training objective? The technique is known as self-conditioning right? Is it an issue with conditional Tweedie's?
-
comment
Comment #44919716
AI with ability but without responsibility is not enough for dramatic socioeconomic change, I think. For now, the critical unique power of human workers is that you can hold them r…
-
comment
Comment #44578257
Has anyone come across any really cool artifacts? I'd be curious to see
-
comment
Comment #44486617
Stablecoins transferred $27 trillion in 2024 - more than Visa and Mastercard combined. This is right in the article. Stablecoins operate using decentralized ledgers on e.g. Ethereu…
-
comment
Comment #44063937
Gemini has beat it already, but using a different and notably more helpful harness. The creator has said they think harness design is the most important factor right now, and that …
-
comment
Comment #43606140
Huh, I imagined this was because of relaxing regulation.
-
comment
Comment #43391115
Good read, thanks for sharing
-
comment
Comment #43287945
> But what is the original purpose of AI research? I will speak for myself here, but I know many other AI researchers will say the same: the ultimate goal is to understand how huma…
-
comment
Comment #43052445
This ought to be called the qwerty effect, for how the qwerty keyboard layout can't be usurped at this point. It was at the right place at the right time, even though arguably its …
-
comment
Comment #42885367
What's surprising about this is how sparsely defined the rewards are. Even if the model learns the formatting reward, if it never chances upon a solution, there isn't any feedback/…
-
comment
Comment #42862418
I used sublime from 2013 to 2021. It was great. Since, I've switched to VS Code and haven't looked back.
-
comment
Comment #42269336
Context is a challenge for LLMs, but the challenge feels of a different quality to me, than the challenge of incorporating local context into automated decision-making AI like algo…
-
comment
Comment #42132502
My interest was piqued, but the extrapolation in [1] is uh... not the most convincing. If there were more data points then sure, maybe
-
comment
Comment #41541148
OK - there's always a nonzero chance of hallucination. There's also a non-zero chance that macroscale objects can do quantum tunnelling, but no one is arguing that we "need to live…