Live data from Hacker News

Viewing profile — shenberg

shenberg

HN member
Joined
Mon, Aug 29, 2011, 7:03 PM UTC
HN karma
409
Public activity
120 items

About shenberg

No profile information was provided.

Recent public activity

  1. comment
    Comment #48966359

    We take a lot of shortcuts when speaking, it's actually much harder to transcribe phonemes than to transcribe words, even when aware of the language being spoken. Some models have …

  2. comment
    Comment #48635665

    You always either go left or down so total 40 steps, choose 20 to be down (or 20 to be right)

  3. comment
    Comment #48537586

    Some example >1B companies off the top of my head: DataDog, Sentry, Snowflake, Okta, MongoDB

  4. comment
    Comment #48458089

    moondream is a beast

  5. comment
    Comment #48442300

    Seems 100% AI generated and automated, the judge also seems suspect - in the first one it's actually GPT-5.5 pro which has the correct email RE: the deepseek one will match a@b.com…

  6. comment
    Comment #48219521

    11% MFU does not mean 89% of GPUs are idle, it means that they're using the GPUs ineffectively.

  7. comment
    Comment #47164728

    Using existing enterprise apps probably - this solution is scalable for the vendor and it's easier to sell using existing software as-is than to start out by writing new custom too…

  8. comment
    Comment #47127645

    Mid-way I realized this was AI writing (took me a while), then I read a quote in the text about a comment that "The tragedy isn’t that they cheated; it’s that the system was design…

  9. comment
    Comment #47033127

    Moshi was an amazing tech demo, building the entire stack from scratch in 6 months with a small team was an amazing show of skill: 7B text LLM data + training, emotive TTS for synt…

  10. comment
    Comment #46898465

    Location: Paris, France (US citizen, EU resident) Remote: Yes Willing to relocate: Not until 2027 Technologies: ML / DS: PyTorch, CUDA, distributed training & inference, performanc…

  11. comment
    Comment #46884929

    There are two ingredients that don't fit in the "attention-is-kernel-smoothing" as far as I can tell: positional encoding and causal masking (another way to say positional encoding…

  12. comment
    Comment #46777318

    I don't understand how using group-theory language to describe number-theoretic properties provides extra insight in this case (e.g. conjecture: all perfect numbers are even is mor…

  13. comment
    Comment #46402983

    ssh exe.dev works

  14. comment
    Comment #46002735

    The short and unsatisfying answer is that an LLM generation is a markov chain, except that instead of counting n-grams in order to generate the posterior distribution, the training…

  15. comment
    Comment #45760748

    When countries like North Korea, which depends on cybercrime to fund itself, are signatories, you have to wonder whether this agreement means what its title says.

  16. comment
    Comment #44708661

    The reality of meetings in most places I've seen is that key stakeholders have already formed an opinion beforehand, the meeting is a place to disseminate decisions that have alrea…

  17. comment
    Comment #44387971

    When I read "51% fewer false positives" followed immediately by "Median comments per pull request cut by half" it makes me wonder how many true positives they find. That's maybe un…

  18. comment
    Comment #42855604

    The DeepSeek v3 model had a net training cost of >$5m for the final training run, the paper lists over 100 authors[1], meaning highly-paid engineers. This is also one of a sequence…

  19. comment
    Comment #42724230

    That's really not true, e.g. the wikipedia page on population transfer in the Ottoman empire[1]. This dates way back to the Assyrian and Persian empries explicitly moving conquered…

  20. comment
    Comment #42203422

    Anecdotally, a pro-audio software company I worked with had to fire 1/3 of the company when their copy-protection was cracked and sales tanked immediately afterwards, and recovered…

  21. comment
    Comment #42034311

    Under the leaderboard tab, if the "Solution" column has an icon, it's clickable. 2nd place solution is by Jeremy Howard (of fast.ai fame), which I'd summarize as TrueSkill Through …

  22. comment
    Comment #40309787

    The CLIP plot (Fig. 2) is damning, however some of the generative models show flat responses in Fig. 3 (e.g. Adobe GigaGAN, DALL-E-mini). While those are on the one hand technicall…

  23. comment
    Comment #39500189

    I would have expected a sham-treatment arm to the experiment, because how do you differentiate between "an intervention 30 minutes beforehand caused improved learning" and "our spe…

  24. comment
    Comment #39287529

    I suspect that weight initializations are geared towards inputs being normal random variables with mean 0 and variance 1. Deviating from that makes the learning process unhappy.

  25. comment