Live data from Hacker News

Viewing profile — gergelycsegzi

gergelycsegzi

HN member
Joined
Tue, Mar 12, 2024, 1:29 AM UTC
HN karma
36
Public activity
22 items

About gergelycsegzi

No profile information was provided.

Recent public activity

  1. comment
    Comment #48760851

    Re 1 - that is a very kind offer! Our current public template library is very limited, so let me come back to you on this. 2. We see exactly the same thing. There is a trade-off in…

  2. comment
    Comment #48757623

    Yes, we do it by having multiple stages to the pipeline. First we would extract the independent data points (from say both page 4 and 40) and a second pass step establishes relatio…

  3. comment
    Comment #48757534

    This does indeed look really interesting. We have deterministic validations (and some deterministic excel transformations) but using more deterministic transformations for text bas…

  4. comment
    Comment #48754382

    If Claude is good enough for your use case then for sure. If you need scale, persistent structure and verifiability we can help:)

  5. comment
    Comment #48753778

    Haha thanks, the reader can try and guess which is which;) We actually don't use embeddings or vector similarity, since those tend not to work well in specialist domains (e.g. for …

  6. comment
    Comment #48753186

    I can see why, it's tempting to go for full automation. The reason we go for fine grained sourcing is so that people can build their awareness quickly. Plus many of our customers w…

  7. comment
    Comment #48752690

    Potentially, but at that scale cost and latency may actually become an issue, so probably better to consider some sort of indexing or keyword searching.

  8. comment
    Comment #48751327

    100% the really hard challenge is that the intermediate representation (ie the parquet equivalent) will be dependent on the given use case. So what we do with the platform is have …

  9. comment
    Comment #48751285

    I'll need to check it out! We had the same observation in that the possible space is almost endless, and for example even for the same file type there may be different kind of proc…

  10. comment
    Comment #48751018

    We were also surprised at first. The reason the models don't do so well is that they need to find information across 90k pages. When they are pointed to the right location they ten…

  11. comment
    Comment #48750490

    Hey, that's exactly it!

  12. comment
    Comment #48750309

    Fully agree, that's why we quite like the Databricks OfficeQA benchmark.. it made us experts on historical US treasuries haha Some screenshots in here: https://www.parsewise.ai/off…

  13. comment
    Comment #48750064

    Similar to my other comment, we assume that llamaparse and others can provide the individual page OCR. But once you have that the way that you can integrate it into your workflows …

  14. comment
    Comment #48749984

    Hey, good point about structure for integrated workflows:) Fully agree, for enterprises we need to guarantee types, flag discrepancies and provide underlying sources so they can in…

  15. comment
    Comment #48749608

    In practice we find that each domain (and even each organisation) ends up having highly customized definitions. At first, fairly generic templated definitions sort of work, but wha…

  16. comment
    Comment #48749021

    Haha no appreciate it! That's on me for not calling it out explicitly (was trying to make the video as short as possible), but the demo UIs were literally vibe coded to show the ea…

  17. comment
    Comment #48748259

    Planning to serve good things for sure, and appreciate your note. Ofc I didn't agree with everything Palantir was doing (also to the extent that we even knew about them at the time…

  18. comment
    Comment #48748076

    Great question! 1. We are working with the assumption that OCR is (or soon will be) solved at super low prices. So if we have the extracted data, what can we do with it? Where we s…

  19. comment
    Comment #48747853

    I learnt a lot at Palantir, though always worked in commercial so no ties to security state (for the better or worse). (Also side-note, we are working towards enabling frontier per…

  20. comment
    Comment #48747731

    "That is a great catch!"

  21. comment
    Comment #48746974

    Ah probably should add a link to our website: https://www.parsewise.ai/api

  22. story
    Launch HN: Parsewise (YC P25) – Reason Across Documents with an API

    Hi all, it’s Greg and Max, founders of Parsewise here ( https://www.parsewise.ai/api ). Parsewise transforms a bucket of unstructured data into schema compliant data, retaining lin…