Copyright and intellectual property are directly at odds with these types of efforts, and has been losing to linux, gnu, github, wikipedia, mit open courseware, youtube, LLMs and their datasets. But copyright did slay Napster, PirateBay, anna’s archive etc…
A multimodal dataset with one trillion tokens
51–55 of 55 posts
Re: A multimodal dataset with one trillion tokens
#52Earlier quoted context omitted.
Until individual countries start realizing that protecting copyright is costing them lots of potential economic growth coming from IA, and the IA business start lobbying more than the copyright business, at which point the law would just change. Intellectual property is a fairly recent invention in economic history, and it only happened because it benefited the elite. If the balance of power changes so will the law.
IA?
Re: A multimodal dataset with one trillion tokens
#53How effective would modeling raw byte sequences be, with the individual bytes as the "tokens", and a vocabulary of 256 elements? You could then train on any kind of digital data.
You might be interested in reading up on DNA sequence llm models and tooling.
Re: A multimodal dataset with one trillion tokens
#54Does it make sense to measure a dataset in tokens? Shouldn't it be tokenizer-agnostic? I.e. the OpenAI tokenizer encodes about ~4 characters per token, but I could also have a tokenizer that does 1 character per token leading to a ~4x increase in token count (relative to the OpenAI tokenizer.)
Hello! Totally agree that tokens will be model dependent. We chose to calculate tokens using the GPT-2 tokenizer as that is a common metric used by other datasets like fineweb. So this should roughly give you a sense of how large the data is in comparison to others. We report other metrics too like number of documents and number of images.
Re: A multimodal dataset with one trillion tokens
#55Earlier quoted context omitted.
This all be quite dated in 10-20 years now. Common information will be free as it was in the 90s, but valuable information will then probably cost even more. And 99.9% times illegal to obtain or possess.
Until individual countries start realizing that protecting copyright is costing them lots of potential economic growth coming from IA, and the IA business start lobbying more than the copyright business, at which point the law would just change. Intellectual property is a fairly recent invention in economic history, and it only happened because it benefited the elite. If the balance of power changes so will the law.