Live data from Hacker News

Confirmed: Reflection 70B's official API is a wrapper for Sonnet 3.5

old.reddit.com

31–40 of 70 posts

Re: Confirmed: Reflection 70B's official API is a wrapper for Sonnet 3.5

#31

Okay, let's think this through step by step. Isn't 'reflection thinking' a pretty well known technique in the AI prompt field? So this model was supposed to be so much better... why, exactly? It makes very little sense to me. Is it just about separating the "reflections/chain of thoughts" from the "final output" via specific tags?

Even though this was a scam, it's somewhat plausible. You finetune on synthetic data with lots of common reasoning mistakes followed by self-correction. You also finetine on synthetic data without reasoning mistakes where the "reflection" says that everything is fine. The model then learns to recognize output with subtle mistakes/hallucinations due to having been trained to do that.

Re: Confirmed: Reflection 70B's official API is a wrapper for Sonnet 3.5

#32
post #22
post #6

Context: someone announced a Llama 3.1 70B fine tune with incredible benchmark results a few days ago. It's been a dramatic ride: - The weight releases were messed up: released Lora for Llama 3.0, claiming it was a 3.1 fine tune - Evals initially didn't meet expectations when run on released weights - The evals starting performing near/at SOTA when using a hosted endpoint - Folks are finding clever ways to see what m…

Who is Sahil Chaudhary? Why he doesn't announce such a great advancement himself? Why Matt Shumer first announces it only because -- according to a later claim on X.com -- he trusted Sahil, does that mean Matt is unable to participate most of the progress? Then why announce a breakthrough without mentioning he was not fully involved to a level he can verify the result in the first place?

As far as I can tell he's the founder of GlaiveAI. There were messages suggesting Matt was an investor, but I haven't been able to confirm this.

Re: Confirmed: Reflection 70B's official API is a wrapper for Sonnet 3.5

#33
post #12
post #6

Context: someone announced a Llama 3.1 70B fine tune with incredible benchmark results a few days ago. It's been a dramatic ride: - The weight releases were messed up: released Lora for Llama 3.0, claiming it was a 3.1 fine tune - Evals initially didn't meet expectations when run on released weights - The evals starting performing near/at SOTA when using a hosted endpoint - Folks are finding clever ways to see what m…

When they were using the Sonnet 3.5 API, they censored the word "Claude" and replaced "Anthropic" with "Meta", then later when people realized this, they removed it. Also, after GPT-4o they switched to a llama checkpoint (probably 405B-inst), so now the tokenizer is in common (no more tokenization trick).

Yeah I managed to get it to admit that it was Claude without much effort (telling it not to lie), and then it magically stopped doing that. FWIW Constitutional AI is great.

Re: Confirmed: Reflection 70B's official API is a wrapper for Sonnet 3.5

#34
post #13

It's amazing what people will do for clout. His whole reputation is ruined. What was Schumer's endgame?

It's also amazing that GlaiveAI will be synonymous with fraud in ML now, because an investor decided to fake some benchmarks. The founder of GlaiveAI, Sahil Chaudhary also participated in the creation of the model.

I wonder if the other investors will sue.

Re: Confirmed: Reflection 70B's official API is a wrapper for Sonnet 3.5

#35

It's amazing what people will do for clout. His whole reputation is ruined. What was Schumer's endgame?

Plenty of people have scammed their way to the top of the benchmark league tables, by training on the benchmarking datasets. And a lot of the people who do this just get ignored - they don't take much heat for it.

If the scam hadn't gained enough publicity for people to start paying attention, he would have gotten away with it :)

Re: Confirmed: Reflection 70B's official API is a wrapper for Sonnet 3.5

#36

It's amazing what people will do for clout. His whole reputation is ruined. What was Schumer's endgame?

But does reputation work? Will people google "Matt Shumer scam", "HyperWrite scam", "OthersideAI scam", "Sahil Chaudhary scam", "Glaive AI scam" before using their products? He wasted everyone's time, but what's the downside for him? Lots of influencers did fraud, and they do just fine.

Re: Confirmed: Reflection 70B's official API is a wrapper for Sonnet 3.5

#38
post #32
post #22

Earlier quoted context omitted.

Who is Sahil Chaudhary? Why he doesn't announce such a great advancement himself? Why Matt Shumer first announces it only because -- according to a later claim on X.com -- he trusted Sahil, does that mean Matt is unable to participate most of the progress? Then why announce a breakthrough without mentioning he was not fully involved to a level he can verify the result in the first place?

As far as I can tell he's the founder of GlaiveAI. There were messages suggesting Matt was an investor, but I haven't been able to confirm this.

Matt said it was approximately ”$1000" and that he has disclosed it "before" in a reply. https://x.com/mattshumer_/status/1832558298509275440
Post reply on HN