Live data from Hacker News

Viewing profile — aesthesia

aesthesia

HN member
Joined
Thu, Sep 06, 2018, 3:24 PM UTC
HN karma
742
Public activity
193 items

About aesthesia

No profile information was provided.

Recent public activity

  1. comment
    Comment #49225795

    What makes you think that this is actually good PR for the firms involved? Every claim that this is good PR comes from someone who has increased their negative views of OpenAI base…

  2. comment
    Comment #49185847

    I'm not sure we should take it as a given that AI will not also become better than humans at mathematical exposition.

  3. comment
    Comment #49119304

    Ah, I see, taking "to test A" as an infinitive. But it would be strange to say test B is an alternative without saying what it's an alternative to. The other reading still seems qu…

  4. comment
    Comment #49119290

    You can make this kind of claim about living in lots of different places. Some amount of engineering and technology has been necessary to enable humans to live in most parts of the…

  5. comment
    Comment #49117998

    In those experiments, the effect only happened between models from the same family (and likely the same weight initialization).

  6. comment
    Comment #49117236

    How is that ambiguous? The best interpretation I can find where "test" is a verb is an elision: Test [that] B is an alternative to test A. That is an unlikely reading: "test" is a …

  7. comment
    Comment #49114752

    Luna's now cheaper than 5.4-nano (for output tokens). That's a significant improvement.

  8. comment
    Comment #49100314

    These aren't assertions being made without arguments or evidence. Maybe you've engaged with those and find them unconvincing. But from what you've said in this thread it seems more…

  9. comment
    Comment #49093972

    Quines produce similar issues to self-referential sentences without being directly self-referential. e.g. "'Yields falsehood when preceded by its quotation' yields falsehood when p…

  10. comment
    Comment #49093102

    I mean, Cervantes was parodying medieval romances. That's what Don Quixote was reading, not modern novels. The point is, it should give you at least some pause when the people most…

  11. comment
    Comment #49092684

    Fun, though as hinted at the end, the point of LLM "truth" probes is to measure the model's internal judgment of truthfulness. There's no reason this judgment, even if measured wit…

  12. comment
    Comment #49092611

    How many of those concerns came from the people building the technologies? John Carmack wasn't up in arms about how Doom was going to destroy society. Charles Dickens didn't think …

  13. comment
    Comment #49089395

    I think the origin of the title formula is Timothy Chow's "You Could Have Invented Spectral Sequences" ( https://www.ams.org/notices/200601/fea-chow.pdf ), which is certainly not a…

  14. comment
    Comment #49077553

    Who pays to build the open weight models? The cost of training a frontier model is orders of magnitude greater than the cost of safety tests.

  15. comment
    Comment #49043375

    Yeah, there's probably not a lot systematic behind Anthropic's version numbers. Opus 4.5 was a third of the price of Opus 4.1, indicating there was probably a change in underlying …

  16. comment
    Comment #49043002

    Unless you believe that OpenAI was trying to get their models to break out of the sandbox and hack into Hugging Face, and aiding them in that goal, I don't see how any of those que…

  17. comment
    Comment #49042919

    Could there, conceivably, be some third factor at play here? One that both led OpenAI to have concerns about releasing GPT-2 as well as led Microsoft to invest in OpenAI?

  18. comment
    Comment #49040522

    They could start by releasing the MAI-1 models they announced recently. https://microsoft.ai/models/

  19. comment
    Comment #49040480

    You ask good questions, and I would like to know the answers to them, but I don't think any likely answers would actually change the conclusion much. Any probable set of circumstan…

  20. comment
    Comment #49040230

    Where do you see the claim that "long-horizon goals in real world settings are now effectively settled"? The argument you put in their mouth would be a bad one, but I don't see any…

  21. comment
    Comment #49039779

    This is a comment about user-facing responses, which are seldom the thing you're worried about when thinking about token efficiency.

  22. comment
    Comment #49029533

    It depends a lot on the model. Some models (e.g. Qwen) hew very closely to the party line even without external guardrails, while others (including Kimi, it seems) are much more ev…

  23. comment
    Comment #49013409

    1. OpenAI being bad at managing risk from misaligned models is not evidence that their models are not misaligned. It's evidence that they're not taking misalignment seriously. 2. H…

  24. comment
    Comment #48999878

    My guess is that RL training being done with particular generation parameters makes models much more brittle to changes in these parameters, and that's why we're seeing changes lik…

  25. comment
    Comment #48998513

    And what do those who encourage its creation get?