Viewing profile — lorey
lorey
HN member- Joined
- Sat, Dec 06, 2014, 1:46 PM UTC
- HN karma
- 636
- Public activity
- 84 items
- HN profile
- View on Hacker News ↗
About lorey
- https://karllorey.com
- https://github.com/lorey
- https://startupradar.co
- https://markets.apistemic.com
Recent public activity
-
comment
Comment #49196044
Found this to display the optimal LLM choice while building evalry. It's such a useful tool, not only for thinking about it, but for visualization, too. Example: Which LLM gives me…
-
comment
Comment #49095403
What I read from that is that there's a chance claude gets worse using a claude.md and that's a usability issue on their side.
-
comment
Comment #48675145
Their response: > The team that made dataroom has stated that they did not use any of papermark’s code and that dataroom was made from scratch with inspiration from existing docume…
-
comment
Comment #46718677
That is not what the article argues.
-
comment
Comment #46718672
Haha, very true. Exactly as described in the article.
-
comment
Comment #46718661
This is true with one caveat. In most cases, e.g. with regular ML, evals are easy and not doing them results in inferior performance. With LLMs, especially frontier LLMs, this has …
-
comment
Comment #46718633
This is a very good point. When I came in, the founder did a lot of evaluation based on a few prompts and with manual evaluation, exactly as described. Showing the results helped m…
-
comment
Comment #46711903
Doesn't this depend a lot on private vs company usage? There's no way I could spend more than a few hundreds alone, but when you run prompts on 1M entities in some corporate use ca…
-
comment
Comment #46709777
It's not you, it's the HN hug of death. There's so much load on the server, I'm barely able to download the redis image I need for caching...
-
comment
Comment #46709721
Thanks. Will take a look.
-
comment
Comment #46709650
Depends on your remaining budget ;)
-
comment
Comment #46709456
I've skipped that in the article, but absolutely!
-
comment
Comment #46709198
Fixed, thanks. Not a native speaker.
-
comment
Comment #46698447
That's interesting. Similarly, we found out that for very simple tasks the older Haiku models are interesting as they're cheaper than the latest Haiku models and often perform equa…
-
comment
Comment #46698317
Pushed a fix. Could you check, please? Any resources you can recommend to properly tackle this going forward?
-
comment
Comment #46698133
Will fix, thanks :)
-
comment
Comment #46698064
Totally agree with your point. While I can't say specifically, it's a traditional (German) business he's doing vertically integrated with AI. Customer support is really bad in this…
-
comment
Comment #46698037
This went straight to prod, even earlier than I'd opted for. What do you mean?
-
comment
Comment #46698017
Appreciate the feedback, will work on that.
-
comment
Comment #46698004
You're right. We did a few use cases and I have to admit that while customer service is easiest to explain, its where I'd also not choose the cheapest model for said reasons.
-
comment
Comment #46697968
Yes, absolutely. This aligns with what we found. It seems to be necessary to be very clear on scoring (at least for Opus 4.5).
- story
-
comment
Comment #46438012
Very interesting points. Would you mind sharing a few examples of when cherry-picking is necessary and why atomic changes are a lie? I'm using a monorepo for my company across 3+ p…
-
comment
Comment #46210195
Yes, there's the an other specific tags providing far more options than the old favicons. For anyone interested: https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/...
-
comment
Comment #46210126
Thanks, this is how I feel about this, too.