Live data from Hacker News

Viewing profile — zone411

zone411

HN member
Joined
Fri, Aug 13, 2010, 12:49 PM UTC
HN karma
4,261
Public activity
939 items

About zone411

https://twitter.com/LechMazur

10 LLM benchmarks: https://github.com/lechmazur/

https://www.linkedin.com/in/lech-mazur-69b70493/

Advameg (City-data.com) founder and CEO. AI startup founder.

Author: AI melody songwriting assistant https://melodies.ai

Author: Accurate COVID-19 county-by-county neural net case prediction model based on most data.

Recent public activity

  1. comment
    Comment #49184921

    A complete mischaracterization, as usual for HN lately when discussing AI or LessWrong. Obviously, even average levels of persuasion are enough to convince some people. And nobody …

  2. comment
    Comment #49044612

    So why don't companies in other industries rush to prove their products are dangerous weapons? Maybe because it would be a really dumb PR stunt?

  3. comment
    Comment #49044610

    And how do people saying this know the capabilities of yet unreleased models?

  4. comment
    Comment #49044603

    How is it in their interest? Scaring customers, worrying employees, and inviting regulators to act is in their interest?

  5. comment
    Comment #49044587

    It's not good marketing for them. This is a talking point with zero evidence that people repeat mindlessly. Scaring customers, worrying employees, and inviting regulators to act wo…

  6. comment
  7. story
  8. comment
    Comment #48544298

    Yes, definitely not a new idea. I had a multi-turn composite model in 2024 that was outperforming the top models across benchmarks: https://x.com/LechMazur/status/18288044850339925…

  9. comment
    Comment #48401633

    That's not proof. Emergent intelligence is not consciousness.

  10. comment
    Comment #48319302

    I’ve tested this model on four of my benchmarks: https://github.com/lechmazur/buyout_game 10th out 36. https://github.com/lechmazur/pact/ 14th out 25. https://github.com/lechmazur/…

  11. comment
    Comment #48283683

    100%. It's sad to see that this attitude has spread to HN

  12. comment
    Comment #48215681

    I actually tried using GPT-5.5 Pro on this problem recently. It thought it was making progress on one path, but it made so many mistakes that it didn't feel worth it pushing furthe…

  13. story
  14. comment
    Comment #47707585

    [flagged]

  15. comment
    Comment #47707539

    https://variety.com/2020/digital/news/twitter-unblocks-new-y...

  16. story
  17. comment
    Comment #47556280

    I built this benchmark this month: https://github.com/lechmazur/sycophancy . There are large differences between LLMs. There are large differences between LLMs. For example, Mistra…

  18. comment
    Comment #47556226

    I built two related benchmarks this month: https://github.com/lechmazur/sycophancy and https://github.com/lechmazur/persuasion . There are large differences between LLMs. For examp…

  19. story
  20. comment
    Comment #47495156

    Hmm, maybe in the next edition, Opus gets expensive. I should probably run GPT-5.4 xhigh too if I do that for fairness...

  21. story
  22. comment
    Comment #47403320

    Rationalists were right about everything that mattered: crypto, AI, COVID... HN commentators, by contrast, were wrong about everything that mattered.

  23. story
  24. comment
    Comment #47267296

    Results from my Extended NYT Connections benchmark: GPT-5.4 extra high scores 94.0 (GPT-5.2 extra high scored 88.6). GPT-5.4 medium scores 92.0 (GPT-5.2 medium scored 71.4). GPT-5.…

  25. comment
    Comment #47157409

    I've made top-10 lists of LLMs' favorite names to use in creative writing here: https://x.com/LechMazur/status/2020206185190945178 . They often recur across different LLMs. For exa…