Earlier quoted context omitted.
What makes you think that. Apple is the company that would be most successful at hiding something like this then introduce it as siri ai or something. Not that they are I am just saying Apple keeps everything close to its chest when it comes to products it might introduce in the future.
I work in the field and they just are not hiring the people they need to be hiring.
Llama 2
591–600 of 860 posts
Re: Llama 2
#592Earlier quoted context omitted.
Is it possible that some LLM’s are trained on these benchmarks? Which would mean they’re overfitting and are incorrectly ranked? Or am I misunderstanding these benchmarks?…
Test leakage is not impossible for some benchmarks. But researchers try to avoid/mitigate that as much as possible for obvious reasons.
Re: Llama 2
#593I asked llama2.ai for some personal advice to see what insights it might offer, it responded: tthtthtthtthtthtth tthtthtthtthtthtth tthtthtthtthtth tthtthtthtthtth tthtthttht tthtthtth tthtth thtth th thtth thtth thtth thtth tth tth tth tthtth tth tth tthtth tthtth tthtth tthtth tthtth ttht tthtth tthtth tthtth tthtth thtthtth thtthtthtth thtthtthtth thtthtth tthtthtth thttht thtthtth thtthtth thtthtth thtth thttht t…
Re: Llama 2
#594Earlier quoted context omitted.
Google's LLMs are all vaporware. No one's ever seen them. They're supposedly mind-blowing but when they are released they always sound like lobotomized monkeys. All the AlphaGo/AlphaFold stuff is very cool, but since no one has seen their LLMs this is about as convincing as my claiming I've donated billions to charity.
I can assure you Google BERT isn't vaporware. It was probably a challenge to integrate it into search, but they did that. So your assertion has been refuted based on your use of "all", at the very least.
Re: Llama 2
#595Llama2 failed pretty hard. "FTP traffic is not typically used for legitimate purposes."
Re: Llama 2
#596Earlier quoted context omitted.
People keep saying this is commoditize your complement but that's not what this is! Goods A and B are economic complements if, when the price of A goes down, demand for B goes up. LLMs are not complements to social media platforms. There is zero evidence that if "the price of LLMs goes down" then "demand for social media apps go up". This is a case of commoditizing the competition but that's not the same thing. Commo…
>LLMs are not complements to social media platforms Tell that to the people generating text for social media campaigns using LLMs.
Re: Llama 2
#597Plugged in a prompt I've been developing for use in a potential product at work (using chatgpt previously). Llama2 failed pretty hard. "FTP traffic is not typically used for legitimate purposes."
Re: Llama 2
#598Earlier quoted context omitted.
Sparse MoE models are neither new nor secret. The only reason you haven't seen much use of them for LLMs is because they would typically well underperform their dense counterparts. Until this paper ( https://arxiv.org/abs/2305.14705 ) indicated they apparently benefit far more from Instruct tuning than dense models, it was mostly a "good on paper" kind of thing. In the paper, you can see the underperformance i'm talk…
This paper came out well after GPT-4, so apparently this was indeed a secret before then.
We also have no indication sparse models outperform dense counterparts so it's scale either way.
Re: Llama 2
#599Earlier quoted context omitted.
Test leakage is not impossible for some benchmarks. But researchers try to avoid/mitigate that as much as possible for obvious reasons.
Given all of the times OpenAI has trained on peoples' examples of "bad" prompts, I am sure they are fine-tuning on these benchmarks. It's the natural thing to do if you are trying to position yourself as the "most accurate" AI.
If it performs about as well in instances it has never seen before (test set) then it's not overfit to the test.