Earlier quoted context omitted.
it allows the program to reference memory without having to manage it in the heap space. it would make the program faster in a memory managed language, otherwise it would reduce the memory footprint consumed by the program.
You mean it converts an expression like `buf[i]` into a baroque sequence of CPU exception paths, potentially involving a trap back into the kernel.
Zpdf: PDF text extraction in Zig
41–50 of 88 posts
Re: Zpdf: PDF text extraction in Zig
#42Earlier quoted context omitted.
fixed.
Yeah, sorry for confusion. When said Unicode, meant foreign text rather (just) the unescaped symbols, e.g. Greek. At one random Greek textbook[0], zpdf output is (extract | head -15): 01F9020101FC020401F9020301FB02070205020800030209020701FF01F90203020901F9012D020A0201020101FF01FB01FE0208 0200012E0219021802160218013202120222 0209021D0212021D012E013202200222000301FA021A0220021C022002160213012E0222000F000301F90206012C 0…
Re: Zpdf: PDF text extraction in Zig
#43Earlier quoted context omitted.
Why would you make Claude write your commit message for a commit you've spent years working on though?
1. Be not good at or a fan of git when coding 2. Be not good at or a fan of git when committing Not sure what the disconnect is. Now if it were vibecoded, I wouldn't be surprised. But benefit of the doubt
I don't particularly care, though, and I'm more positive about LLMs than negative even if I don't (yet?) use them very much. I think it's hilarious that a few people asked for Python bindings and then bam, done, and one person is like "..wha?" Yes, LLMs can do that sort of grunt work now! How cool, if kind of pointless. Couldn't the cycles have just been spent on trying to make muPDF better? Though I see they're in C and AGPL, I suppose either is motivation enough to do a rewrite instead. (This is MIT Licensed though it's still unclear to me how 100% or even large-% vibe-coded code deserves any copyright protection, I think all such should generally be under the Unlicense/public domain.)
If the intent of "benefit of the doubt" is to reduce people having a freak out over anyone who dares use these tools, I get that.
Re: Zpdf: PDF text extraction in Zig
#44Earlier quoted context omitted.
1. Be not good at or a fan of git when coding 2. Be not good at or a fan of git when committing Not sure what the disconnect is. Now if it were vibecoded, I wouldn't be surprised. But benefit of the doubt
We're well beyond benefit of the doubt these days. If it looks like a duck... For me there wasn't any doubt, the author's first top comment here was evidence enough, then seeing the readme + random code + random commit message, it's all obvious LLM-speak to me. I don't particularly care, though, and I'm more positive about LLMs than negative even if I don't (yet?) use them very much. I think it's hilarious that a few…
I'll try my best to make it a really good one!
Re: Zpdf: PDF text extraction in Zig
#45Earlier quoted context omitted.
You mean it converts an expression like `buf[i]` into a baroque sequence of CPU exception paths, potentially involving a trap back into the kernel.
I don't fully understand the under the hood mechanics of mmap, but I can sense that you're trying to convey that mmap shouldn't be used a blanket optimization technique as there are tradeoffs in terms of page fault overheads (being at the mercy of OS page cache mechanics)
Re: Zpdf: PDF text extraction in Zig
#46Re: Zpdf: PDF text extraction in Zig
#47excellent stuff what makes zig so fast
Not being slow - they compile straight to bytecode, they aren't interpreted, and have aggressive, opinionated optimizations baked in by default, so it's even faster than compiled c (under default conditions.) Contrasted with python, which is interpreted, has a clunky runtime, minimal optimizations, and all sorts of choices that result in slow, redundant, and also slow, performance. The price for performance is safety…
machine code, not https://en.wikipedia.org/wiki/Bytecode
> The price for performance is safety checks
In Zig, non-ReleaseFast build modes have significant safety checks.
> luajit ... with better-than-c performance
No.
Re: Zpdf: PDF text extraction in Zig
#48I built a PDF text extraction library in Zig that's significantly faster than MuPDF for text extraction workloads. ~41K pages/sec peak throughput. Key choices: memory-mapped I/O, SIMD string search, parallel page extraction, streaming output. Handles CID fonts, incremental updates, all common compression filters. ~5,000 lines, no dependencies, compiles in Why it's fast: - Memory-mapped file I/O (no read syscalls) - Z…
a better speed comparison would either be multi-process pdfium (since pdfium was forked from foxit before multi-thread support, you can't thread it), multi-threaded foxit, or something like syncfusion (which is quite fast and supports multiple threads). Or even single thread pdfium vs single thread your-code.
These were always the fastest/best options. I can (and do) achieve 41k pages/sec or better on these options.
The other thing it doesn't appear you mention is whether you handle putting the words in reading order (IE how they appear on the page), or only stream order (which varies in its relation to apperance order) .
If it's only stream order, sure, that's really fast to do. But also not anywhere near as helpful as reading order, which is what other text-extraction engines do.
Looking at the code, it looks like the code to do reading order exists, but is not what is being benchmarked or used by default?
If so, this is really comparing apples and oranges.
Re: Zpdf: PDF text extraction in Zig
#49Earlier quoted context omitted.
You mean it converts an expression like `buf[i]` into a baroque sequence of CPU exception paths, potentially involving a trap back into the kernel.
I don't fully understand the under the hood mechanics of mmap, but I can sense that you're trying to convey that mmap shouldn't be used a blanket optimization technique as there are tradeoffs in terms of page fault overheads (being at the mercy of OS page cache mechanics)
Re: Zpdf: PDF text extraction in Zig
#50I'm curious about the trade-offs mentioned in the comments regarding Unicode handling. For document analysis pipelines (like extracting text from technical documentation or research papers), robust Unicode support is often critical.
Would be interesting to see benchmarks on different PDF types - academic papers with equations, scanned documents with OCR layers, and complex layouts with tables. Performance can vary wildly depending on the document structure.