Live data from Hacker News

Viewing profile — MrCheeze

MrCheeze

HN member
Joined
Sat, Jul 15, 2023, 2:11 AM UTC
HN karma
147
Public activity
28 items

About MrCheeze

I did this stuff: https://mrcheeze.github.io/

Recent public activity

  1. comment
    Comment #49077275

    That is in fact Anthropic's entire reason for existence, the belief that AGI is too dangerous to be controlled by OpenAI/Sam Altman. It naturally follows that it would also be too …

  2. comment
    Comment #48845545

    I mean, surely the main motivation for "use an LLM to rewrite a huge project in a new language" was excitement about the shiny new tech that made it possible.

  3. comment
    Comment #48689703

    TBF, they did it first with ada/babbage/curie/davinci. "Sol" is a much weaker branding, though.

  4. comment
    Comment #47076637

    Does anyone understand why LLMs have gotten so good at this? Their ability to generate accurate SVG shapes seems to greatly outshine what I would expect, given their mediocre spati…

  5. comment
    Comment #47055972

    The Claude Plays Pokemon stream with a minimal harness is a far more significant test of model intelligence compared to the Gemini Plays Pokemon stream (which automatically maintai…

  6. comment
    Comment #47054718

    Notably 45 out of the 50 days of improvement were in two specific dungeons (Silph Co and Cinnabar Mansion) where 4.5 was entirely inadequate and was looping the same mistaken ideas…

  7. comment
    Comment #47052536

    In my experience with the models (watching Claude play Pokemon), the models are similar in intelligence, but are very different in how they approach problems: Opus 4.5 hyperfocuses…

  8. comment
    Comment #46343225

    This writeup on the underground puzzle is worth reading, it's a pretty baffling "puzzle" design. https://pokemow.com/Gen2/ShutterPuzzle/ That said, it's definitely Gem's fault that…

  9. comment
    Comment #46343218

    There were no such writeups, 99% of the discussion about difficulties in Crystal were in twitch and discord chats where Google doesn't scrape. (It hadn't yet gotten the public atte…

  10. comment
    Comment #46343190

    It's hard to say for sure because Gemini 3 was only tested with this prompt. But for Gemini 2.5, which is who the prompt was originally written for, yes this does cut down on bad a…

  11. comment
    Comment #45656062

    Exactly what I was going to post. Optimizations like loop unrolling slow down the N64 because keeping the code size small is the most important factor. I think even compilers of th…

  12. comment
    Comment #45019744

    The busy beaver function is interesting precisely because you _can't_ come up with any computable function that grows faster.

  13. comment
    Comment #44895884

    Claude almost universally reacts to everything with a positive exclamation as its first sentence, regardless of whether it's good or bad. If you don't believe me, just watch https:…

  14. comment
    Comment #43900935

    With 2, the real problem is that approximately 0% of the OpenAI employees actually believed in the mission. Pretty much every single one of them signed the letter to the board dema…

  15. comment
    Comment #43242438

    Are you thinking of Bismuth's "Speedrunning as a gateway to scientific endeavours", perhaps? https://www.youtube.com/watch?v=w8_1lQ2KH50

  16. comment
    Comment #43241555

    As an n=1 data point, that was my exact situation for a while. Also a lot of the people who put out high effort stuff are college students, which works for the same reason. More in…

  17. comment
    Comment #43234831

    I've wondered myself why there's so little overlap between these two closely related interests of mine. Some of it seems to be the "But I don't want to cure cancer. I want to turn …

  18. comment
    Comment #42777217

    How long until we get to the point where models know that LLMs get this wrong, and that it is an LLM, and therefore answers wrong on purpose? Has this already happened? (I doubt it…

  19. comment
    Comment #42367780

    In conversational language, "All my hats..." implies that the speaker has at least two hats, which theoretically means that the sentence could be a lie from them having exactly one…

  20. comment
    Comment #42253831

    https://xcancel.com/legit_rumors/status/1861448164084978157#...

  21. comment
    Comment #40862998

    https://www.scottaaronson.com/writings/bignumbers.html

  22. comment
    Comment #40402309

    Two reasons why I don't particularly believe him: 1) Altman's companies have had similar clauses before: https://news.ycombinator.com/item?id=40396787 2) The entire OpenAI board de…

  23. comment
    Comment #39912848

    If, like me, you are suddenly curious what would happen if you added a small fourth body: https://youtu.be/WrahPSY9pf0

  24. comment
    Comment #39607964

    I have just been informed that my above comment is false, the CLIP-L is in fact referring to OpenAI's, despite that also being the name of an OpenCLIP model.

  25. comment
    Comment #39604489

    One of the diagrams says they're using CLIP-G/14 and CLIP-L/14, which are the names of two OpenCLIP models - meaning they're not using OpenAI's CLIP.