Live data from Hacker News

Viewing profile — jneagu

jneagu

HN member
Joined
Fri, May 22, 2020, 3:45 AM UTC
HN karma
3
Public activity
17 items

About jneagu

No profile information was provided.

Recent public activity

  1. story
    Show HN: HalluMix – A Benchmark for Real-World LLM Hallucination Detection

    We built HalluMix to evaluate how well hallucination detectors perform in the kinds of messy, high-stakes environments where LLMs are actually deployed: long-form outputs, multi-do…

  2. comment
    Comment #42730727

    Hey @echollama! Can you share more about what kind of integrations you are looking for? Also, when you say "enable customers to configure their usage" - what types of configuration…

  3. comment
    Comment #41844768

    "trained on licensed content" - this is a bit misleading. The correct framing is in the article: "made with content owner permission." Most (all?) LLMs out there are trained on lic…

  4. story
  5. comment
    Comment #41838117

    I can save you a bit of market research and tell you that’s unfortunately not the case yet in the market today. There are a few reasons for it - the main one in my opinion being th…

  6. comment
    Comment #41838061

    The challenge with A/B experiments is how you design them to have sufficient power and draw a meaningful conclusion out of them. So, you either need a big % difference between the …

  7. story
  8. comment
    Comment #41830868

    Nice! Overall, frameworks that re-prompt an LLM with feedback or failure modes of the original outputs do very well.

  9. comment
    Comment #41829892

    To your latter point - that’s where I think most of the value of LLMs in education is. They can explain code beyond the educational content that’s already available out there. They…

  10. comment
    Comment #41829749

    I am very curious to see how this is going to impact STEM education. Such a big part of an engineer's education happens informally by asking peers, teachers, and strangers question…

  11. comment
    Comment #41829259

    Anecdotally, synthetic data can get good if the generation involves a nugget of human labels/feedback that gets scaled up w/ a generative process.

  12. comment
    Comment #41829234

    Fair point - I actually had parsed OP's sentence differently. I'll edit my comment. I agree, LLMs performance for coding tasks is super biased in favor of well-represented language…

  13. comment
    Comment #41829212

    Yeah, There was a reference in a paywalled article a year ago ( https://www.theinformation.com/articles/openai-made-an-ai-br... ): "Sutskever's breakthrough allowed OpenAI to overc…

  14. comment
    Comment #41829080

    Edit: OP had actually qualified their statement to refer to only underrepresented coding languages. That's 100% true - LLM coding performance is super biased in favor of well-repre…

  15. story
    Show HN: Smell – A framework for aligning LLM evaluators to human feedback

    We've built SMELL (Subject-Matter Expert Language Liaison), a new framework that combines human expertise with LLMs to create feedback-informed, domain-specific LLM evaluators. One…

  16. comment
    Comment #39992719

    Hey HN, We are excited to reveal Quotient, a platform enabling developers to evaluate, improve and ship high-quality AI products through fast, real-world, data-backed experimentati…

  17. story