Viewing profile — shenberg
shenberg
HN member- Joined
- Mon, Aug 29, 2011, 7:03 PM UTC
- HN karma
- 409
- Public activity
- 120 items
- HN profile
- View on Hacker News ↗
About shenberg
No profile information was provided.
Recent public activity
-
comment
Comment #48966359
We take a lot of shortcuts when speaking, it's actually much harder to transcribe phonemes than to transcribe words, even when aware of the language being spoken. Some models have …
-
comment
Comment #48635665
You always either go left or down so total 40 steps, choose 20 to be down (or 20 to be right)
-
comment
Comment #48537586
Some example >1B companies off the top of my head: DataDog, Sentry, Snowflake, Okta, MongoDB
-
comment
Comment #48458089
moondream is a beast
-
comment
Comment #48442300
Seems 100% AI generated and automated, the judge also seems suspect - in the first one it's actually GPT-5.5 pro which has the correct email RE: the deepseek one will match a@b.com…
-
comment
Comment #48219521
11% MFU does not mean 89% of GPUs are idle, it means that they're using the GPUs ineffectively.
-
comment
Comment #47164728
Using existing enterprise apps probably - this solution is scalable for the vendor and it's easier to sell using existing software as-is than to start out by writing new custom too…
-
comment
Comment #47127645
Mid-way I realized this was AI writing (took me a while), then I read a quote in the text about a comment that "The tragedy isn’t that they cheated; it’s that the system was design…
-
comment
Comment #47033127
Moshi was an amazing tech demo, building the entire stack from scratch in 6 months with a small team was an amazing show of skill: 7B text LLM data + training, emotive TTS for synt…
-
comment
Comment #46898465
Location: Paris, France (US citizen, EU resident) Remote: Yes Willing to relocate: Not until 2027 Technologies: ML / DS: PyTorch, CUDA, distributed training & inference, performanc…
-
comment
Comment #46884929
There are two ingredients that don't fit in the "attention-is-kernel-smoothing" as far as I can tell: positional encoding and causal masking (another way to say positional encoding…
-
comment
Comment #46777318
I don't understand how using group-theory language to describe number-theoretic properties provides extra insight in this case (e.g. conjecture: all perfect numbers are even is mor…
-
comment
Comment #46402983
ssh exe.dev works
-
comment
Comment #46002735
The short and unsatisfying answer is that an LLM generation is a markov chain, except that instead of counting n-grams in order to generate the posterior distribution, the training…
-
comment
Comment #45760748
When countries like North Korea, which depends on cybercrime to fund itself, are signatories, you have to wonder whether this agreement means what its title says.
-
comment
Comment #44708661
The reality of meetings in most places I've seen is that key stakeholders have already formed an opinion beforehand, the meeting is a place to disseminate decisions that have alrea…
-
comment
Comment #44387971
When I read "51% fewer false positives" followed immediately by "Median comments per pull request cut by half" it makes me wonder how many true positives they find. That's maybe un…
-
comment
Comment #42855604
The DeepSeek v3 model had a net training cost of >$5m for the final training run, the paper lists over 100 authors[1], meaning highly-paid engineers. This is also one of a sequence…
-
comment
Comment #42724230
That's really not true, e.g. the wikipedia page on population transfer in the Ottoman empire[1]. This dates way back to the Assyrian and Persian empries explicitly moving conquered…
-
comment
Comment #42203422
Anecdotally, a pro-audio software company I worked with had to fire 1/3 of the company when their copy-protection was cracked and sales tanked immediately afterwards, and recovered…
-
comment
Comment #42034311
Under the leaderboard tab, if the "Solution" column has an icon, it's clickable. 2nd place solution is by Jeremy Howard (of fast.ai fame), which I'd summarize as TrueSkill Through …
-
comment
Comment #40309787
The CLIP plot (Fig. 2) is damning, however some of the generative models show flat responses in Fig. 3 (e.g. Adobe GigaGAN, DALL-E-mini). While those are on the one hand technicall…
-
comment
Comment #39500189
I would have expected a sham-treatment arm to the experiment, because how do you differentiate between "an intervention 30 minutes beforehand caused improved learning" and "our spe…
-
comment
Comment #39287529
I suspect that weight initializations are geared towards inputs being normal random variables with mean 0 and variance 1. Deviating from that makes the learning process unhappy.
- comment