Viewing profile — jneagu
jneagu
HN member- Joined
- Fri, May 22, 2020, 3:45 AM UTC
- HN karma
- 3
- Public activity
- 17 items
- HN profile
- View on Hacker News ↗
About jneagu
No profile information was provided.
Recent public activity
-
story
Show HN: HalluMix – A Benchmark for Real-World LLM Hallucination Detection
We built HalluMix to evaluate how well hallucination detectors perform in the kinds of messy, high-stakes environments where LLMs are actually deployed: long-form outputs, multi-do…
-
comment
Comment #42730727
Hey @echollama! Can you share more about what kind of integrations you are looking for? Also, when you say "enable customers to configure their usage" - what types of configuration…
-
comment
Comment #41844768
"trained on licensed content" - this is a bit misleading. The correct framing is in the article: "made with content owner permission." Most (all?) LLMs out there are trained on lic…
- story
-
comment
Comment #41838117
I can save you a bit of market research and tell you that’s unfortunately not the case yet in the market today. There are a few reasons for it - the main one in my opinion being th…
-
comment
Comment #41838061
The challenge with A/B experiments is how you design them to have sufficient power and draw a meaningful conclusion out of them. So, you either need a big % difference between the …
- story
-
comment
Comment #41830868
Nice! Overall, frameworks that re-prompt an LLM with feedback or failure modes of the original outputs do very well.
-
comment
Comment #41829892
To your latter point - that’s where I think most of the value of LLMs in education is. They can explain code beyond the educational content that’s already available out there. They…
-
comment
Comment #41829749
I am very curious to see how this is going to impact STEM education. Such a big part of an engineer's education happens informally by asking peers, teachers, and strangers question…
-
comment
Comment #41829259
Anecdotally, synthetic data can get good if the generation involves a nugget of human labels/feedback that gets scaled up w/ a generative process.
-
comment
Comment #41829234
Fair point - I actually had parsed OP's sentence differently. I'll edit my comment. I agree, LLMs performance for coding tasks is super biased in favor of well-represented language…
-
comment
Comment #41829212
Yeah, There was a reference in a paywalled article a year ago ( https://www.theinformation.com/articles/openai-made-an-ai-br... ): "Sutskever's breakthrough allowed OpenAI to overc…
-
comment
Comment #41829080
Edit: OP had actually qualified their statement to refer to only underrepresented coding languages. That's 100% true - LLM coding performance is super biased in favor of well-repre…
-
story
Show HN: Smell – A framework for aligning LLM evaluators to human feedback
We've built SMELL (Subject-Matter Expert Language Liaison), a new framework that combines human expertise with LLMs to create feedback-informed, domain-specific LLM evaluators. One…
-
comment
Comment #39992719
Hey HN, We are excited to reveal Quotient, a platform enabling developers to evaluate, improve and ship high-quality AI products through fast, real-world, data-backed experimentati…
- story