This will be an interesting test to see how fast you can bootstrap GPT-4 level performance with unlimited funds and talent that already has deep knowledge of the internals. With the initial adoption of ChatGPT alongside Copilot, OpenAI's data moat of crawled data & RLHF is pretty vast. And that's not leaving the walled garden of OpenAI. You can simulate a lot of this using other off-the-shelf LLMs (see Alpaca) but no…
Many of OpenAIs most talented people left to start Anthropic. They have billions in funding and have not yet got particularly close to GPT-4. I think that illustrates it will be a be a big uphill battle for any new entrant no matter how well funded or resourced.
Wrong. Claude 2 beats GPT-4 is some benchmarks (e.g. HumanEval Python coding; math; analytical writing.). It's close enough. It doesn't matter who holds the crown this week, Anthropic definitely has ingredients to make GPT-4-class model.
This is like comparing similar cars from BMW and Toyota, finding few specific parameters where BMW has a higher score and saying "You see? Toyota engineering is nowhere close".
This actually shows Sam Altman's true contribution: the free version of ChatGPT is undeniably worse than Bing Chat, and yet ChatGPT is a bigger brand.
(And it might be a deliberate choice to save money for Claude 3 instead instead of making Claude 2 absolutely SotA.)