Viewing profile — gergelycsegzi
gergelycsegzi
HN member- Joined
- Tue, Mar 12, 2024, 1:29 AM UTC
- HN karma
- 36
- Public activity
- 22 items
- HN profile
- View on Hacker News ↗
About gergelycsegzi
No profile information was provided.
Recent public activity
-
comment
Comment #48760851
Re 1 - that is a very kind offer! Our current public template library is very limited, so let me come back to you on this. 2. We see exactly the same thing. There is a trade-off in…
-
comment
Comment #48757623
Yes, we do it by having multiple stages to the pipeline. First we would extract the independent data points (from say both page 4 and 40) and a second pass step establishes relatio…
-
comment
Comment #48757534
This does indeed look really interesting. We have deterministic validations (and some deterministic excel transformations) but using more deterministic transformations for text bas…
-
comment
Comment #48754382
If Claude is good enough for your use case then for sure. If you need scale, persistent structure and verifiability we can help:)
-
comment
Comment #48753778
Haha thanks, the reader can try and guess which is which;) We actually don't use embeddings or vector similarity, since those tend not to work well in specialist domains (e.g. for …
-
comment
Comment #48753186
I can see why, it's tempting to go for full automation. The reason we go for fine grained sourcing is so that people can build their awareness quickly. Plus many of our customers w…
-
comment
Comment #48752690
Potentially, but at that scale cost and latency may actually become an issue, so probably better to consider some sort of indexing or keyword searching.
-
comment
Comment #48751327
100% the really hard challenge is that the intermediate representation (ie the parquet equivalent) will be dependent on the given use case. So what we do with the platform is have …
-
comment
Comment #48751285
I'll need to check it out! We had the same observation in that the possible space is almost endless, and for example even for the same file type there may be different kind of proc…
-
comment
Comment #48751018
We were also surprised at first. The reason the models don't do so well is that they need to find information across 90k pages. When they are pointed to the right location they ten…
-
comment
Comment #48750490
Hey, that's exactly it!
-
comment
Comment #48750309
Fully agree, that's why we quite like the Databricks OfficeQA benchmark.. it made us experts on historical US treasuries haha Some screenshots in here: https://www.parsewise.ai/off…
-
comment
Comment #48750064
Similar to my other comment, we assume that llamaparse and others can provide the individual page OCR. But once you have that the way that you can integrate it into your workflows …
-
comment
Comment #48749984
Hey, good point about structure for integrated workflows:) Fully agree, for enterprises we need to guarantee types, flag discrepancies and provide underlying sources so they can in…
-
comment
Comment #48749608
In practice we find that each domain (and even each organisation) ends up having highly customized definitions. At first, fairly generic templated definitions sort of work, but wha…
-
comment
Comment #48749021
Haha no appreciate it! That's on me for not calling it out explicitly (was trying to make the video as short as possible), but the demo UIs were literally vibe coded to show the ea…
-
comment
Comment #48748259
Planning to serve good things for sure, and appreciate your note. Ofc I didn't agree with everything Palantir was doing (also to the extent that we even knew about them at the time…
-
comment
Comment #48748076
Great question! 1. We are working with the assumption that OCR is (or soon will be) solved at super low prices. So if we have the extracted data, what can we do with it? Where we s…
-
comment
Comment #48747853
I learnt a lot at Palantir, though always worked in commercial so no ties to security state (for the better or worse). (Also side-note, we are working towards enabling frontier per…
-
comment
Comment #48747731
"That is a great catch!"
-
comment
Comment #48746974
Ah probably should add a link to our website: https://www.parsewise.ai/api
-
story
Launch HN: Parsewise (YC P25) – Reason Across Documents with an API
Hi all, it’s Greg and Max, founders of Parsewise here ( https://www.parsewise.ai/api ). Parsewise transforms a bucket of unstructured data into schema compliant data, retaining lin…