Viewing profile — yelmahallawy
yelmahallawy
HN member- Joined
- Tue, Apr 13, 2021, 10:06 PM UTC
- HN karma
- 8
- Public activity
- 33 items
- HN profile
- View on Hacker News ↗
About yelmahallawy
Recent public activity
-
comment
Comment #47360521
Gotchu. Yeah that's pretty quick, awesome thanks!
-
comment
Comment #47360509
Makes sense, simpler=better. Thanks!
-
comment
Comment #47347694
The guy who reviews all of this, is his role in the company fully dedicated to reviewing these eval pipelines?
-
comment
Comment #47347678
Yeah, it feels like an unsolved problem still. I've also seen many teams spend hours on human review in eval pipelines (and this accumulates with each new model that gets released)…
-
comment
Comment #47347659
Have some good takeaways / feedback on this? First time I hear about Braintrust (the eval platform) so I'll look into it but I'm curious on your experience with it so far.
-
comment
Comment #47347642
And I think this is a common problem actually — figuring out what to measure and how to measure it – it's not black and white. What I do is have a few dimensions to measure it agai…
-
comment
Comment #47347627
I'd love to hear more about what you're working on (if you're open to sharing!). I like to play with knowledge base powered chatbots but what's most useful to me (and probably my p…
-
comment
Comment #47347575
Any takeaways on Promptfoo?
-
comment
Comment #47347566
Any takeaways? Has it been helpful? OpenAI just acquired them so it's probably useful but I was curious to hear more from people who've actually used it.
-
comment
Comment #47347558
Yeah it's a super tedious process and I was hoping that _maybe_ there is a tool out there that can help with this.
-
comment
Comment #47347540
Ah good read, thanks for sharing!
-
comment
Comment #47347519
Yeah that's essentially what I'm looking for. Since now that AI has become such a core part of most businesses, it's pretty critical to use the _best_ models + prompts for whatever…
-
comment
Comment #47347510
Ah, interesting – yeah only swapping out the model isn't super insightful since models perform differently given different prompts. I'm going to look into GEPA, thanks!
-
comment
Comment #47347481
This is interesting approach, thanks for the insight! If I may ask, _approximately_ how long does it take to test a newly-released model with the current strategy?
-
comment
Comment #47347439
Do you play with the temperature/top k parameters at all?
-
comment
Comment #47347429
This makes sense. I am particularly interested in your invoice processing app example because the accuracy of those outputs can be quantitatively measured from 0%-100% accuracy. I'…
-
comment
Comment #47347361
What I've noticed is that it's hard to measure outputs that aren't binary right or wrong, and that's where most human intervention is needed. The biggest examples of this are chatb…
-
story
Ask HN: How are people doing AI evals these days?
With the buzz that's happening with all the new AI models that get released (what feels like every other week), how are companies running internal AI evals to determine which model…
- story
-
comment
Comment #37762080
Location: Toronto/San Francisco Bay area Remote: Yes Willing to relocate: Yes Technologies: React, Next.js, Node.js, Typescript, PostgreSQL, GraphQL, Resume: https://drive.google.c…
- story
-
story
Show HN: Track and bill users' AI token usage through an API
Nowadays, everyone's building an AI app, but they're not pricing their apps fairly. Each user should be billed separately based on their GPT token usage, and many developers aren't…
-
comment
Comment #35825793
By having users pay for credits. There's 3 different types of credits, prompt credits, chat credits, and image generation credits. Prompt credits are used when generate prompts, ch…
-
comment
Comment #35824079
I just built a website that uses GPT-4 to create prompts for ChatGPT, DALLE•2, and Midjourney where you can test out the prompts directly on the website. I'm looking for some feedb…
- story