Live data from Hacker News

Viewing profile — eli

eli

HN member
Joined
Fri, Feb 23, 2007, 4:21 AM UTC
HN karma
30,067
Public activity
9,905 items

About eli

Co-Founder of Industry Dive. We are a digital media company that publishes business news and original analysis for executives in different industries. industrydive.com

Some examples:

https://www.utilitydive.com/ - News for electric utility industry

https://www.biopharmadive.com/ - News for Biotech & Pharma industry

https://www.supplychaindive.com/ - Supply chain & logistics news

more: https://www.industrydive.com/industries/

Email: eli.at.elidickinson.com

Web: http://www.elidickinson.com/

Recent public activity

  1. comment
    Comment #49215109

    Qwen 3.8 Max is very strong at troubleshooting and code review.

  2. comment
    Comment #49214778

    I assume/hope this is about prices going up for the next release of Pro

  3. comment
    Comment #49213858

    And all that just to allow internet access for npm and pypi? If you've got the bandwidth and disk space, it's very easy to make an offline mirror of both.

  4. comment
    Comment #49202066

    That's got a significant selection bias. Claude and ChatGPT and Gemini and other subs do not go through openrouter.

  5. comment
    Comment #49201258

    I use https://pi.dev/ which works fine out of the box but is fairly minimal and intended to be customized. There are many extensions. OpenCode or oh-my-pi might make more sense if …

  6. comment
    Comment #49201233

    It's not exactly what you asked for, but the public transit and multimodal focused apps are generally much better at this stuff. I like CityMapper and Transit. Neither is good for …

  7. comment
    Comment #49201040

    It's not enough that it's better? Many providers will host it and will compete on price. It also can't easily be taken away because one company (or one government) decides they don…

  8. comment
    Comment #49201026

    I believe it. It's extremely good at troubleshooting. I gave Qwen and Kimi K3 the same annoying, complicated, intermittent bug to track down. Kimi did a bit better in understanding…

  9. comment
    Comment #49186346

    The amount of abuse of free VMs must be absolutely immense even with credit card verification. I’m not sure there needs to be additional secret motives.

  10. comment
    Comment #49173526

    They're usually leased. So yeah they probably owe Flock new ones.

  11. comment
    Comment #49173389

    They bill out the ass to lease them

  12. comment
    Comment #49173368

    I'm not familiar with North Dakota in particular, but at the federal level and in some states ALL works created by the government are exempt from copyright. The public paid for it,…

  13. comment
    Comment #49159249

    According to a few tasks from my little personal coding benchmark it's very good at coding and kinda bad at web design. (Also excellent at "draw me a picture" one-shot prompts, for…

  14. comment
    Comment #49157531

    To be fair, "people with no skills inflating their own value" is what LinkedIn has always been like. But I guess LLMs are uniquely well positioned for that task.

  15. comment
    Comment #49117133

    These are the API rates. They don’t train on prompts from the API. In what sense are you the product?

  16. comment
    Comment #49117058

    I mean, yeah. Alternatively I’ve seen the approach of starter with the cheaper/faster model but give it a tool to handoff or talk to a strong model if it gets stuck.

  17. comment
    Comment #49111106

    This is the improved iOS parental controls. When they shipped it (and for years afterward) it simply did not work https://www.wsj.com/tech/personal-tech/apples-parental-contr... It…

  18. comment
    Comment #49044185

    Yup, trying to be really strategic about testing. I didn't end up sticking with it, but I tried requiring test cases to cite a matching clause in the task assignment. But also: the…

  19. comment
    Comment #49042541

    Max is lot of extra reasoning. I wonder how many fewer tasks it solves on high. I bet that costs quite a lot less.

  20. comment
    Comment #49042523

    Oh sorry I meant the framework. Though I could see posting a couple of the tasks.

  21. comment
    Comment #49042371

    I have actually been planning to open source the framework. Only benchmarks I care about are the ones that look like real work I do. So it makes it easy to trawl your own repos loo…

  22. comment
    Comment #49042210

    In my tests, it averages to much cheaper than Opus 4.8 on real tasks on account of being smarter and more token efficient. I have a benchmark to build a game engine from a set of w…

  23. comment
    Comment #49042027

    It's a new and improved version of an existing model? I don't think it's intentionally befuddling.

  24. comment
    Comment #49042017

    That's a plausible explanation but I'm not seeing evidence for it. I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and F…

  25. comment
    Comment #49029805

    I’m not convinced it wasn’t a PR stunt from the start