Live data from Hacker News

Viewing profile — raunakchowdhuri

raunakchowdhuri

HN member
Joined
Tue, Nov 22, 2022, 10:07 PM UTC
HN karma
99
Public activity
34 items

About raunakchowdhuri

No profile information was provided.

Recent public activity

  1. comment
    Comment #47669037

    The big one is that LLMs get lazy on repetitive tasks. They'll skip rows or consolidate entries instead of grinding through every last one. So you need verify-and-re-extract loops …

  2. comment
    Comment #47668975

    We've made a lot of changes in the past few months that make our standard extract much, much better, as well as Deep Extract for documents even longer than that. We'd love for you …

  3. story
  4. story
  5. comment
    Comment #46036234

    Have a slack channel with them, these are the versions they mentioned: posthog-node 4.18.1 posthog-js 1.297.3 posthog-react-native 4.11.1 posthog-docusaurus 2.0.6

  6. story
  7. comment
    Comment #44361869

    We're fixing it! This for some reason happens on only _some_ phones in our office so was hard to repro. I think has to do with Safari rendering. Will tone down our WebGPU usage

  8. comment
    Comment #44360374

    dang no way! we were both in boston too

  9. comment
    Comment #44358729

    this is exactly where we're going with this! glad you see the vision :)

  10. comment
  11. comment
  12. story
  13. comment
    Comment #43288111

    comparisons to more outputs coming soon!

  14. comment
    Comment #43287278

    We ran some benchmarks comparing against Gemini Flash 2.0. You can find the full writeup here: https://reducto.ai/blog/lvm-ocr-accuracy-mistral-gemini A high level summary is that …

  15. story
  16. comment
    Comment #42954289

    CTO of Reducto here. Love this writeup! We’ve generally found that Gemini 2.0 is a great model and have tested this (and nearly every VLM) very extensively. A big part of our resea…

  17. comment
    Comment #42953853

    would encourage you to take a look at some of the real data here! https://huggingface.co/spaces/reducto/rd_table_bench you'll find that most of the errors here are structural issue…

  18. comment
    Comment #42054322

    Love the Pubtables work! It's a really useful dataset. Their data comes from existing annotations from scientific papers, so in our experience it doesn't include a lot of the harde…

  19. story
    Rd-TableBench – Accurately evaluating table extraction

    Hey HN! A ton of document parsing solutions have been coming out lately, each claiming SOTA with little evidence. A lot of these turned out to be LLM or LVM wrappers that hallucina…

  20. comment
    Comment #41358144

    hmmm idk how I would feel about giving an llm cluster access from a security pov

  21. comment
    Comment #40237255

    Interesting... how did you do the scraping of the documentation?

  22. story
    Show HN: Reducto – A vision based document ingestion API for LLMs

    Hey HN, I'm Raunak from Reducto ( https://reducto.ai ), a high-quality document ingestion API tailored for language models. We developed Reducto to address our own need - no existi…

  23. comment
    Comment #38027519

    Hey HN, A few weeks ago I shared a beta of Remembrall here and got a lot of great feedback from people in this community, with one of the most common requests being to open source …

  24. story
  25. comment
    Comment #37532225

    Yep - full export will always be supported for your data. No customer # restriction. I have some optimizations in the works such that the secondary gpt 3.5 call only gets triggered…