Live data from Hacker News

Viewing profile — guilamu

guilamu

HN member
Joined
Wed, Oct 02, 2013, 4:26 PM UTC
HN karma
1,004
Public activity
265 items

About guilamu

No profile information was provided.

Recent public activity

  1. story
  2. comment
    Comment #49215923

    Where does this "10x" comes from?

  3. comment
    Comment #48959696

    FTFY: Starting July 20 Claude Fable 5 will ve excluded from Pro plans.

  4. comment
    Comment #48794339

    https://phosh.mobi/faq/#so-what-phones-are-supported

  5. comment
    Comment #48757132

    Wouldn't you say that Valve is an exception to that rule?

  6. story
  7. comment
    Comment #48450491

    Any source to backup this claim, pretty please?

  8. story
  9. comment
    Comment #48248897

    If you're concerned about that do not give internet to your tv and use any kind of tv box instead (shield tv, apple tv, etc).

  10. comment
    Comment #48172228

    It's not. France: €0.149/kWh (~$0.175) US: ~$0.12–$0.14/kWh https://www.globalpetrolprices.com/France/electricity_prices...

  11. comment
    Comment #47919664

    You're right, I've certainly been a bit presumptuous to call this'a benchmark'. It is indeed a flawed test. Yet,It's been giving me the occasion to try some open source models and …

  12. comment
    Comment #47919567

    Most people, including me, beg to disagree. Better Call Saul was a masterpiece. https://www.metacritic.com/tv/better-call-saul/

  13. comment
    Comment #47896428

    Yeah as I said this a benchmark for my usecase only, a single use case, which is obvisouly not representative of everybody's needs. What strike me as very strange though is that 0 …

  14. comment
    Comment #47895498

    https://openrouter.ai/openai/gpt-5.5-pro 30/180 usd on Openrouter. Did I miss something?

  15. comment
    Comment #47895428

    When nothing is noted it's max reasoning (xhigh in copilot chat in vscode if available). The models not availble on copilot were tested through opencode (max reasoning) and deepsee…

  16. comment
    Comment #47895363

    Yes those two models were tested on my own PC (local inference using my own CPU/GPU). So something my be bugged on my setup. gemma4-26b should be far better than gemma4-e4b.

  17. comment
    Comment #47895353

    Yes, the prompt is slim by design. I might be wrong, but the point was to see what the model can do "on it's own". The eval prompt is quite extensive: https://github.com/guilamu/ll…

  18. comment
    Comment #47895322

    Haha, just fixed the date! I haven't evaluated the judge benchmark. You have everything needed in the repo to do so though, so be my guest. It took me a bit of time to put all this…

  19. comment
    Comment #47895291

    Yes Opus 4.7 fast (no reasoning) did a worst job than Sonnet 4.6 high (with reasoning) according to Gemini 3.1 Pro evaluation.

  20. comment
    Comment #47895122

    Just tested it on my homemade Wordpress+GravityForms benchmark and it's one of the worst model of the leaderboard performance wise and the worst value wise: https://github.com/guil…

  21. comment
    Comment #47886545

    Good point! I think I won't change anything right now or I'll have to remake all tests... I'll use your input for the Level 2 task I plan on working on.

  22. story
    Show HN: I blind-tested 14 LLMs on a WP plugin task. Surprising Findings

    Recently, GitHub Copilot silently dropped support for Claude Opus on Pro accounts. Since Opus was my go-to model for my daily workflow (developing WordPress plugins), I needed a re…

  23. comment
    Comment #47839160

    Indeed. I'm still sad thought :(

  24. story
  25. story