Viewing profile — guilamu
guilamu
HN member- Joined
- Wed, Oct 02, 2013, 4:26 PM UTC
- HN karma
- 1,004
- Public activity
- 265 items
- HN profile
- View on Hacker News ↗
About guilamu
No profile information was provided.
Recent public activity
- story
-
comment
Comment #49215923
Where does this "10x" comes from?
-
comment
Comment #48959696
FTFY: Starting July 20 Claude Fable 5 will ve excluded from Pro plans.
-
comment
Comment #48794339
https://phosh.mobi/faq/#so-what-phones-are-supported
-
comment
Comment #48757132
Wouldn't you say that Valve is an exception to that rule?
- story
-
comment
Comment #48450491
Any source to backup this claim, pretty please?
- story
-
comment
Comment #48248897
If you're concerned about that do not give internet to your tv and use any kind of tv box instead (shield tv, apple tv, etc).
-
comment
Comment #48172228
It's not. France: €0.149/kWh (~$0.175) US: ~$0.12–$0.14/kWh https://www.globalpetrolprices.com/France/electricity_prices...
-
comment
Comment #47919664
You're right, I've certainly been a bit presumptuous to call this'a benchmark'. It is indeed a flawed test. Yet,It's been giving me the occasion to try some open source models and …
-
comment
Comment #47919567
Most people, including me, beg to disagree. Better Call Saul was a masterpiece. https://www.metacritic.com/tv/better-call-saul/
-
comment
Comment #47896428
Yeah as I said this a benchmark for my usecase only, a single use case, which is obvisouly not representative of everybody's needs. What strike me as very strange though is that 0 …
-
comment
Comment #47895498
https://openrouter.ai/openai/gpt-5.5-pro 30/180 usd on Openrouter. Did I miss something?
-
comment
Comment #47895428
When nothing is noted it's max reasoning (xhigh in copilot chat in vscode if available). The models not availble on copilot were tested through opencode (max reasoning) and deepsee…
-
comment
Comment #47895363
Yes those two models were tested on my own PC (local inference using my own CPU/GPU). So something my be bugged on my setup. gemma4-26b should be far better than gemma4-e4b.
-
comment
Comment #47895353
Yes, the prompt is slim by design. I might be wrong, but the point was to see what the model can do "on it's own". The eval prompt is quite extensive: https://github.com/guilamu/ll…
-
comment
Comment #47895322
Haha, just fixed the date! I haven't evaluated the judge benchmark. You have everything needed in the repo to do so though, so be my guest. It took me a bit of time to put all this…
-
comment
Comment #47895291
Yes Opus 4.7 fast (no reasoning) did a worst job than Sonnet 4.6 high (with reasoning) according to Gemini 3.1 Pro evaluation.
-
comment
Comment #47895122
Just tested it on my homemade Wordpress+GravityForms benchmark and it's one of the worst model of the leaderboard performance wise and the worst value wise: https://github.com/guil…
-
comment
Comment #47886545
Good point! I think I won't change anything right now or I'll have to remake all tests... I'll use your input for the Level 2 task I plan on working on.
-
story
Show HN: I blind-tested 14 LLMs on a WP plugin task. Surprising Findings
Recently, GitHub Copilot silently dropped support for Claude Opus on Pro accounts. Since Opus was my go-to model for my daily workflow (developing WordPress plugins), I needed a re…
-
comment
Comment #47839160
Indeed. I'm still sad thought :(
- story
- story