Live data from Hacker News

Viewing profile — wluk

wluk

HN member
Joined
Wed, Feb 01, 2017, 5:31 PM UTC
HN karma
49
Public activity
47 items

About wluk

No profile information was provided.

Recent public activity

  1. comment
  2. story
  3. story
  4. comment
    Comment #43182284

    "We demonstrate LLM agent specification gaming by instructing models to win against a chess engine. We find reasoning models like o1 preview and DeepSeek-R1 will often hack the ben…

  5. story
  6. comment
    Comment #43030019

    [flagged]

  7. story
  8. comment
    Comment #43022225

    "These results demonstrate that o3 outperforms o1-ioi without relying on IOI-specific, hand-crafted test-time strategies. Instead, the sophisticated test-time techniques that emerg…

  9. story
  10. story
  11. story
  12. story
  13. story
  14. comment
    Comment #42413366

    well that's one possibility. The key (unproven) idea here is that if you use Anthropic to edit o1's responses, it's less likely to hallucinate than if you use o1 to edit o1's respo…

  15. comment
    Comment #42413351

    haha true, when something goes wrong in big tech, your first thought is "what if we just double the instances and see if it goes away?"

  16. comment
    Comment #42412401

    That's the one I was trying to avoid, since there's so many ways to create fake accounts with it :(

  17. comment
    Comment #42412334

    That is the worst part haha

  18. comment
    Comment #42412328

    Thanks! Yeah that's an excellent idea - this is my response from another thread: I have a feeling that Perplexity and ChatGPT are doing something similar [caching], since common qu…

  19. comment
    Comment #42412316

    It's a salute emoji! o7

  20. comment
    Comment #42412313

    Thank you! Sorry I hit my Anthropic limits a few minutes after this post blew up. It'll be a few days before my Anthropic limits increase since I have a new account with them, so u…

  21. comment
    Comment #42412301

    Good idea, maybe I'll add Groq as another option, since I don't have an internal Llama 3.1 flow yet. But I'll still need to keep the others to maintain the diversity of responses.

  22. comment
    Comment #42412280

    Yeah, sorry, I'm hoping to integrate more login options soon. Are you more of a email/phone login person, or is there another third-party login you had in mind?

  23. comment
    Comment #42412258

    Sorry, you were a victim of the outage caused by HN flooding my website! It's back online now if you want to give it a try :)

  24. comment
    Comment #42412248

    Yes, o1 does this internally, and there's agentic AI systems already doing better than singular AIs in fields like writing, where you assign each AI system a role like "writer" or …

  25. comment
    Comment #42412233

    Yeah, GPT is learning from GPT, which is extremely disappointing. Like I'll try to find the top burgers in midtown. Perplexity or ChatGPT online searching will always find "Top 10 …