Viewing profile — pama
pama
HN member- Joined
- Mon, Oct 11, 2010, 4:08 AM UTC
- HN karma
- 3,696
- Public activity
- 1,163 items
- HN profile
- View on Hacker News ↗
About pama
No profile information was provided.
Recent public activity
-
comment
Comment #49236552
Alzheimers has a very long development stage compared to almost all other dementia (Lewis body also has a long timeline, but none of the other forms has anywhere close to the known…
-
comment
Comment #49215510
No secrets—all published. Very efficient attention. Excellent kernels. Great caching subsystem. Small and well trained model.
-
comment
Comment #49195171
Google’s trajectory unfortunately starts to feel similar to when yahoo lost everything in the early dotcom bubble except it owned a huge chunk of Alibaba. Google’s internal decisio…
-
comment
Comment #49193770
Large groups suffer communication bottlenecks, so by Amdahl’s law will only be as fast as a smaller group. You can of course have many independent small groups, but this is trivial…
-
comment
Comment #49187287
My point is exactly that we cannot even begin to start the process that will lead to “know everything” there is to know unless we make it a science first. Ad hoc statistical models…
-
comment
Comment #49181986
The early physics background is messy and incorrect. I didnt read the full position paper, but from its start: The Lorentz transformations were by Lorentz, well before Einstein’s p…
-
comment
Comment #49181699
Agreed in principle for today’s new ones, though with all the buying frenzy it quickly becomes hard to procure anything but the latest hardware and the economics of older generatio…
-
comment
Comment #49181376
The new ones will be though. Look up the specs for the NVIDIA Vera Rubin: no fans; recycled liquid cooling.
-
comment
Comment #49179226
The paper suggests the opposite of your first statement. The benchmarks become useless because the successive models keep saturating them.
-
comment
Comment #49179208
I agree these are central components, but to avoid oversimplification and the mistaken belief that modern LLMs do a lot of search during inference: If it was so simple, the traditi…
-
comment
Comment #49179109
I wish we knew everything about how LLMs work! We only know very basic elements related to their construction and traning dynamics, and pretty much every major question we would li…
-
comment
Comment #49170278
I agree boring problems exist; bounds may have a fare share of them. None of the bounds problems in this set are even close to this category; many of them are closer to the type of…
-
comment
Comment #49167587
Use nvidia hardware instead and use a larger cluster serving many more users concurrently. Easily 10x–20x higher token rate per GPU with public solutions like dynamo and sglang.
-
comment
Comment #49164531
Your inverted logic does not hold. The fact that such useless problems for bounds exist does not mean that improving bounds is useless. 9 fields medals in the last twenty years, in…
-
comment
Comment #49160506
It is unfair to dismiss contributions to decades old open problems as equivalent to calculating more digits of pi. It missed the mark by a lot—as does the two bucket simplifaction.…
-
comment
Comment #49159316
> Whilst current models can't 'intuit' and come up with conjectures I disagree. I routinely let LLMs speculate or generate hypotheses along the way of helping with technical resear…
-
comment
Comment #49159098
Not sure what your first sentence means, or why you are quoting AI. Many of these problems individually were math at a level approaching the highest possible for expert human mathe…
-
comment
Comment #49156955
It is much more subtle than your specific example, which is strictly speaking a bug, though ofc it has been used during pretraining for efficiency purposes. Sglang and miles have b…
-
comment
Comment #49155173
Any US organization can rent compute.
-
comment
Comment #49152864
You can do whatever you want with the model within your own organization. If you use it commercially—either as a model-as-a-service business or in a very large-scale product—you sh…
-
comment
Comment #49142411
Or just fix it. I use lockdown mode on iphone. It does not work on safari or chrome, ironically telling me to use a different browser like safari or chrome.
-
comment
Comment #49101633
Serious question: if one were willing to give up on curses, isn’t Emacs/elisp providing the best multiplexing system available to humankind? And conceptually, why would agents ever…
-
comment
Comment #49058171
The drug candidates that enter human clinical trials, (phase 1 to phase 3) are identical in all ways to the final product if approved. The company is not allowed to change any part…
-
comment
Comment #49045163
What is the rationale for gpt-5.5 when gpt-5.6-sol exists?
-
comment
Comment #49039777
The internet convention in the 90s used to be two dashes for the en-dash character and three dashes for em-dash character. You used a space-dash-space notation that is easy to deci…