A multimodal dataset with one trillion tokens
31–40 of 55 posts
Re: A multimodal dataset with one trillion tokens
#32Earlier quoted context omitted.
Salesforce has long been involved in publishing quality NLP papers, especially during Stephen Merity's tenure. Smerity's papers are some of my favourite. Check out https://ar5iv.labs.arxiv.org/html/1708.02182 And my all-time favourite https://ar5iv.labs.arxiv.org/html/1911.11423
Hey thanks for those Smerity links, hadn't run across his work yet, second one in particular looks great
How could you go wrong with a paper that starts with
> Language has been a thorn in humanity’s side since we evolved a complex enough audio and graphics processing unit to grunt, let alone write cryptocurrency whitepapers or opinion columns.
Re: A multimodal dataset with one trillion tokens
#33I havent trained any LLMs, so please accept my comment with all the naivety with which it is given - but in the "examples of MINT multimodal documents" graphic at the top of the README, it feels to me as though the labeling (for the images on the left) couldn't be much worse? Is this normal for these datasets? How are we able to build such powerful models with such poor quality data?
I think most commonly image datasets like this consist of images and their captions, with the presumption that the content author had _some_ reason of associating the two. The goal of the model is to learn that association. And with a _lot_ of examples, to learn nuanced representations.
In the third image, for example, we see some kind of text on a material. The caption mentions "Every year he rides for someone we know, touched by cancer". Perhaps the model is fed another example of bicycle races, with similar imagery of racing bibs. Perhaps its fed another of a race that specifically mentions it's a charity ride to raise money for cancer. Perhaps....
You get the idea. Alone, each example provides only vague connections between the image and the caption. But when you have a ton of data it becomes easier to separate noise from a weak signal.
Re: A multimodal dataset with one trillion tokens
#34Earlier quoted context omitted.
The people building CRM software aren't also the ones doing AI research. The two have nothing to do with each other.
The ones doing AI research would be working at more prestigious institutions.
Re: A multimodal dataset with one trillion tokens
#35I havent trained any LLMs, so please accept my comment with all the naivety with which it is given - but in the "examples of MINT multimodal documents" graphic at the top of the README, it feels to me as though the labeling (for the images on the left) couldn't be much worse? Is this normal for these datasets? How are we able to build such powerful models with such poor quality data?
Not to say that data quality does not matter, but these noisy sets are still very useful.
Re: A multimodal dataset with one trillion tokens
#36Means we can’t legally use it?
Re: A multimodal dataset with one trillion tokens
#37Earlier quoted context omitted.
I’m skeptical of the caliber of talent at Salesforce given the unusable state of their core product.
The people building CRM software aren't also the ones doing AI research. The two have nothing to do with each other.
Salesforce is anything but a CRM software company nowadays.
Re: A multimodal dataset with one trillion tokens
#38They use Bazel and shit, which is an acid test for being professionals, it’s a real shop.
The Magnificent 7 are about to get the taste slapped out of their mouth by skittish momentum guys and their chattels on Sand Hill Road. I look forward to the space this week will create for shops like Salesforce.
Re: A multimodal dataset with one trillion tokens
#39What would the use-case be for this model? What are the advantages over something like Llama?
Re: A multimodal dataset with one trillion tokens
#40So, I read the blog post and checked the Github page, but not a clear picture here for me. I am still kinda new to the LLM space. What would the use-case be for this model? What are the advantages over something like Llama?