Viewing profile — zone411
zone411
HN member- Joined
- Fri, Aug 13, 2010, 12:49 PM UTC
- HN karma
- 4,261
- Public activity
- 939 items
- HN profile
- View on Hacker News ↗
About zone411
10 LLM benchmarks: https://github.com/lechmazur/
https://www.linkedin.com/in/lech-mazur-69b70493/
Advameg (City-data.com) founder and CEO. AI startup founder.
Author: AI melody songwriting assistant https://melodies.ai
Author: Accurate COVID-19 county-by-county neural net case prediction model based on most data.
Recent public activity
-
comment
Comment #49184921
A complete mischaracterization, as usual for HN lately when discussing AI or LessWrong. Obviously, even average levels of persuasion are enough to convince some people. And nobody …
-
comment
Comment #49044612
So why don't companies in other industries rush to prove their products are dangerous weapons? Maybe because it would be a really dumb PR stunt?
-
comment
Comment #49044610
And how do people saying this know the capabilities of yet unreleased models?
-
comment
Comment #49044603
How is it in their interest? Scaring customers, worrying employees, and inviting regulators to act is in their interest?
-
comment
Comment #49044587
It's not good marketing for them. This is a talking point with zero evidence that people repeat mindlessly. Scaring customers, worrying employees, and inviting regulators to act wo…
-
comment
Comment #48996872
[dead]
- story
-
comment
Comment #48544298
Yes, definitely not a new idea. I had a multi-turn composite model in 2024 that was outperforming the top models across benchmarks: https://x.com/LechMazur/status/18288044850339925…
-
comment
Comment #48401633
That's not proof. Emergent intelligence is not consciousness.
-
comment
Comment #48319302
I’ve tested this model on four of my benchmarks: https://github.com/lechmazur/buyout_game 10th out 36. https://github.com/lechmazur/pact/ 14th out 25. https://github.com/lechmazur/…
-
comment
Comment #48283683
100%. It's sad to see that this attitude has spread to HN
-
comment
Comment #48215681
I actually tried using GPT-5.5 Pro on this problem recently. It thought it was making progress on one path, but it made so many mistakes that it didn't feel worth it pushing furthe…
- story
-
comment
Comment #47707585
[flagged]
-
comment
Comment #47707539
https://variety.com/2020/digital/news/twitter-unblocks-new-y...
- story
-
comment
Comment #47556280
I built this benchmark this month: https://github.com/lechmazur/sycophancy . There are large differences between LLMs. There are large differences between LLMs. For example, Mistra…
-
comment
Comment #47556226
I built two related benchmarks this month: https://github.com/lechmazur/sycophancy and https://github.com/lechmazur/persuasion . There are large differences between LLMs. For examp…
- story
-
comment
Comment #47495156
Hmm, maybe in the next edition, Opus gets expensive. I should probably run GPT-5.4 xhigh too if I do that for fairness...
- story
-
comment
Comment #47403320
Rationalists were right about everything that mattered: crypto, AI, COVID... HN commentators, by contrast, were wrong about everything that mattered.
- story
-
comment
Comment #47267296
Results from my Extended NYT Connections benchmark: GPT-5.4 extra high scores 94.0 (GPT-5.2 extra high scored 88.6). GPT-5.4 medium scores 92.0 (GPT-5.2 medium scored 71.4). GPT-5.…
-
comment
Comment #47157409
I've made top-10 lists of LLMs' favorite names to use in creative writing here: https://x.com/LechMazur/status/2020206185190945178 . They often recur across different LLMs. For exa…