Live data from Hacker News

GigaToken: ~1000x faster Language model tokenization

github.com

11–20 of 137 posts

Re: GigaToken: ~1000x faster Language model tokenization

#11
post #2

What sort of setups do people have that are bounded by the speed of the tokenizer?

Author here! In my case it's mostly pretraining experiments, where you might want to change your data mixture/filtering/processing of training data, and splits are usually done at a token-level instead of a text level. In this case we usually run for days on a huge number of CPUs to finish tokenizing something like DCLM. From what I can tell it's also useful for inference when considering time-to-first-token (TTFT) a…

Very cool, thanks.

Re: GigaToken: ~1000x faster Language model tokenization

#12
post #2

What sort of setups do people have that are bounded by the speed of the tokenizer?

Pre-training data is pre-tokenized ahead of time before being used to not waste any GPU compute.

A massive speedup like this is a nice efficiency savings on some of these data pipelines for sure.

Re: GigaToken: ~1000x faster Language model tokenization

#15
post #10
post #6

Earlier quoted context omitted.

Wait, since when does it matter whether something being hyper-optimized is useful? The computer going brrrr on an interesting problem is in itself the goal!

That's fair, I just figure there are useful scenarios as well. Apologies if I came off as dismissive!

It didn’t come off as dismissive to me. I was curious as well as to where such optimizing helps and knew that the answers to your question would help me discover use cases I didn’t think of

Re: GigaToken: ~1000x faster Language model tokenization

#16
post #2

What sort of setups do people have that are bounded by the speed of the tokenizer?

If you are training an LLM, you need to tokenize the text before it’s trained on. A lot of time this can be done in parallel with the GPU though.

I have spent way too much time waiting 10-15 minutes tokenizing my training dataset only for the run to crash over some minor bug after that. (If I was smarter, I’d test on a smaller batch first.)

Re: GigaToken: ~1000x faster Language model tokenization

#18
post #2

What sort of setups do people have that are bounded by the speed of the tokenizer?

I've data where i cannot store metadata that i need to search semantically so i embed it on the fly at every search with static embedding and tokenizing was more than 99% of the cpu time. Granted that was due the naive implementation of the default tokenizer which was o^2 with document length and just switching to a proper scanner solved most of it without going to simd and whatnot, but still.

Re: GigaToken: ~1000x faster Language model tokenization

#20
post #9

Earlier quoted context omitted.

Author here! In my case it's mostly pretraining experiments, where you might want to change your data mixture/filtering/processing of training data, and splits are usually done at a token-level instead of a text level. In this case we usually run for days on a huge number of CPUs to finish tokenizing something like DCLM. From what I can tell it's also useful for inference when considering time-to-first-token (TTFT) a…

Can't you tokenize in preloading on demand?

You can, but this usually results in sequences with padding/truncation, since you won't know how many tokens your inputs map to before you actually tokenize them. This also makes shuffling difficult.

In practice every training project I've worked on does tokenization in a separate data processing phase.

Post reply on HN