Live data from Hacker News

Viewing profile — breadislove

breadislove

HN member
Joined
Thu, Mar 14, 2024, 10:58 PM UTC
HN karma
211
Public activity
70 items

About breadislove

No profile information was provided.

Recent public activity

  1. comment
    Comment #49203239

    we have not converged at all, if you look at how different the chinese models in terms of architecture you can guess that the labs are experimenting a lot as well. we are seeing al…

  2. comment
    Comment #49188008

    On what do you guys test the model. Its very dubious that there is no common retrieval benchmark such as browsecomp plus or similar tested. And what metric do you report?

  3. story
  4. comment
    Comment #48852004

    adam, i'd like to get in touch and would love to run the benachmark with mixedbread as a search backend. we are doing this right now with a lot of compliance companies. would be ve…

  5. comment
    Comment #48769190

    yes, your are right. what heading would you have taken here?

  6. comment
    Comment #48769178

    everything worth writing, you should write yourself

  7. comment
    Comment #48763500

    ah whoops, I'll fix it. ty!

  8. comment
    Comment #48763299

    The ndcg loss is minimal 90.26 -> 89.65. This means it maintains most of the quality.

  9. comment
    Comment #48763212

    to which email did you send it? can u send it to support please?

  10. comment
    Comment #48763197

    this is the reason why we report ndcg and not recall. ndcg respects fine grained details so you get the an overview of how much details you are trading off since it would hurt the …

  11. comment
    Comment #48763178

    yes exactly.

  12. story
  13. comment
    Comment #48588944

    slop complaining about other slop

  14. story
  15. comment
    Comment #48378616

    very bad take. with most modern multomodal models you get way better performance then going to text first

  16. comment
    Comment #46426382

    this might be interesting: https://www.theinformation.com/articles/chatgpt-doctors-star... > $150M RR on just ads, +3x from August. On source: https://x.com/ArfurRock/status/199961…

  17. comment
    Comment #46426374

    a good system (like openevidence) indexes every paper released and semantic search can incredible helpful since the the search api of all those providers are extremely limited in t…

  18. story
  19. story
  20. story
  21. comment
    Comment #45991790

    you should check mixedbread out. we support indexing multimodal data and making data ready for ai. we are adding video and audio support by the end of the year. might be interestin…

  22. story
  23. story
  24. comment
    Comment #45644937

    or deberta but nevertheless super interesting!

  25. comment
    Comment #45643006

    For everyone wondering how good this and other benchmarks are: - the OmniAI benchmark is bad - Instead check OmniDocBench[1] out - Mistral OCR is far far behind most Open Source OC…