Live data from Hacker News

Viewing profile — simonhughes22

simonhughes22

HN member
Joined
Fri, Apr 12, 2013, 3:09 PM UTC
HN karma
186
Public activity
155 items

About simonhughes22

No profile information was provided.

Recent public activity

  1. story
  2. comment
    Comment #38408446

    This is like calling the nuclear anti-proliferationists useless when they had just gotten started. AGI has only been in the general public's consciousness for about a year since Ch…

  3. comment
    Comment #38385315

    Llama 2 (and variants). Has the lowest hallucination rate ( https://github.com/vectara/hallucination-leaderboard ), and its open source and so we know what went into it, and the co…

  4. comment
    Comment #38383945

    Yeah it's odd they chose Palm 2 to compare against. Not a very strong model by most measurements.

  5. comment
    Comment #38383632

    Like a number of other LLMs we tested, including the Palm 2 chat model (chat-bison-001), it adds in the street value, and assumes the plants are cannabis (which is reasonable but i…

  6. comment
    Comment #38383561

    Prompt: You are a chat bot answering questions using data. You must stick to the answers provided solely by the text in the passage provided. You are asked the question 'Provide a …

  7. comment
    Comment #38383556

    The model is bad at hallucinating despite their claims. See the first prompt i tried here: https://twitter.com/hughes_meister/status/172740068973816258...

  8. comment
    Comment #38383453

    This is just typical of so much work in the field. They pick and choose which models to compare against and on which benchmarks. If this model was truly great, they would be compar…

  9. comment
    Comment #38169650

    You can view the responses here in the linked csv file: https://github.com/vectara/hallucination-leaderboard

  10. comment
    Comment #38169633

    The original data we used was not annotated with sources, only where the overall data came from. Most was news articles. The length doesn't seem to matter too much as we see a lot …

  11. comment
    Comment #38169480

    We may write a research paper at some point. For now, see here: https://vectara.com/cut-the-bull-detecting-hallucinations-in... Given the number of models involved, we have over 9k…

  12. comment
    Comment #38169388

    Yes. Just because the model is smaller doesn't always mean by default it's worse, as they may be trained for less time or on less data, which in some cases could be beneficial. The…

  13. comment
    Comment #38169368

    Yes thanks for fixing that.

  14. comment
    Comment #38167648

    I worked on the model with our research team. Recently featured in this NYT ( https://www.nytimes.com/2023/11/06/technology/chatbots-hallu... . Post here to AMA. We are also lookin…

  15. comment
    Comment #38163925

    That's the term used by the academic literature also, so Hallucinate is an industry standard term.

  16. story
  17. comment
    Comment #37792919

    It's not that simple. Originally OpenAI released a model to try and detect whether some content was generated by an LLM or not. They later dropped the service as it wasn't accurate…

  18. comment
    Comment #37665999

    Wondering how many people are now downloading this and other libs like Dart and trying to do stock market prediction or crypto price forecasting. Most of the devs i know, myself in…

  19. comment
    Comment #31260384

    This is why i moved to data science so i can focus more on solving problems than picking frameworks and libraries. We are not completely immune to this problem, but by and large th…

  20. comment
    Comment #30702095

    Thanks that is definitely broken. Logging the issue right now.

  21. comment
    Comment #30700080

    I'd suggest using the app over the mobile website. If you are in store, it will tell you where in the store the items are (Bay and Aisle) which is super useful.

  22. comment
    Comment #30700063

    It does semantic matching. We don't have a lot of exact matches for that search for birch wood (given the exact dimensions) so the engine broadens the search criteria automatically…

  23. comment
    Comment #30700012

    I worked on the system. It's a similar idea but it's on e-commerce products and not websites. So you can't use things like page rank when doing product search.

  24. comment
    Comment #30699775

    The results here look fine to me in positions 3 and 4 - https://www.homedepot.com/s/2'x4'%2520piece%2520of%2520birch... What are you expecting to see? I worked on this search engin…

  25. comment