Live data from Hacker News

A multimodal dataset with one trillion tokens

github.com

51–55 of 55 posts

Re: A multimodal dataset with one trillion tokens

#51

Copyright and intellectual property are directly at odds with these types of efforts, and has been losing to linux, gnu, github, wikipedia, mit open courseware, youtube, LLMs and their datasets. But copyright did slay Napster, PirateBay, anna’s archive etc…

Pirate Bay is thriving.

Re: A multimodal dataset with one trillion tokens

#52
post #46

Earlier quoted context omitted.

Until individual countries start realizing that protecting copyright is costing them lots of potential economic growth coming from IA, and the IA business start lobbying more than the copyright business, at which point the law would just change. Intellectual property is a fairly recent invention in economic history, and it only happened because it benefited the elite. If the balance of power changes so will the law.

IA?

Intellect Amplifier?

Re: A multimodal dataset with one trillion tokens

#53

How effective would modeling raw byte sequences be, with the individual bytes as the "tokens", and a vocabulary of 256 elements? You could then train on any kind of digital data.

You might be interested in reading up on DNA sequence llm models and tooling.

Can you reccomend any resources for this?

Re: A multimodal dataset with one trillion tokens

#54
post #20

Does it make sense to measure a dataset in tokens? Shouldn't it be tokenizer-agnostic? I.e. the OpenAI tokenizer encodes about ~4 characters per token, but I could also have a tokenizer that does 1 character per token leading to a ~4x increase in token count (relative to the OpenAI tokenizer.)

Hello! Totally agree that tokens will be model dependent. We chose to calculate tokens using the GPT-2 tokenizer as that is a common metric used by other datasets like fineweb. So this should roughly give you a sense of how large the data is in comparison to others. We report other metrics too like number of documents and number of images.

How does the GPT-2 tokenizer deal with non-text input? This dataset is multimodal but I thought GPT-2 was text only.

Re: A multimodal dataset with one trillion tokens

#55
post #44

Earlier quoted context omitted.

This all be quite dated in 10-20 years now. Common information will be free as it was in the 90s, but valuable information will then probably cost even more. And 99.9% times illegal to obtain or possess.

Until individual countries start realizing that protecting copyright is costing them lots of potential economic growth coming from IA, and the IA business start lobbying more than the copyright business, at which point the law would just change. Intellectual property is a fairly recent invention in economic history, and it only happened because it benefited the elite. If the balance of power changes so will the law.

thats apparent. my point being that IA will put the final nail in the coffin as it's retelling information in a way which evades copyright in many cases.
Post reply on HN