Live data from Hacker News

DeepMind and OpenAI win gold at ICPC

codeforces.com

161–170 of 255 posts

Re: DeepMind and OpenAI win gold at ICPC

#161

Earlier quoted context omitted.

Firstly, automobiles are really impressive. Second, with that out the way, these cars are not playing the same game as horses… first, and quite obviously they have massive amounts of horsepower, which is kind of like giving a team of horses… many more horses. But also cars have an absolutely massive fuel capacity. Petrol is such an efficient store of chemical energy compared to hay and cars can store gallons of it. I…

Yeah man, and it would be wild to publish an article titled "Ford Mustang and Honda Civic win gold in the 100 meter dash at the Olympics" if what happened was the companies drove their cars 100 meters and tweeted that they did it faster than the Olympians had run. Actually that's too generous, because the humans are given a time limit in ICPC, and there's no clear mapping to say how the LLM's compute should be limite…

Cars going faster than humans or horses isn't very interesting these days, but it was 100+ years ago when cars were first coming on the scene.

We are at that point now with AI, so a more fitting headline analogy would be "In a world first, automobile finishes with gold-winning time in horse race".

Headlines like those were a sign that cars would eventually replace horses in most use-cases, so the fact that we could be in the the same place now with AI and humans is a big deal.

Re: DeepMind and OpenAI win gold at ICPC

#162
post #27

Earlier quoted context omitted.

Why the AI hate? How is it different from sharing your knowledge with another individual or writing a book to share it? > AI companies are not paying anyone for that piece of information So? For the vast majority of human existence, paying for content was not a thing, just like paying for air isn't. The copyright model you are used to may just be too forced. Many countries have no moral qualms about "pirating" Window…

These vigorously held and loudly proclaimed opinions don't matter. Don't waste the mental energy. They're more interested in performative ignorance and argument than anything productive. It's somewhere between trying to engage Luddites during the industrial revolution and having a reasonable discussion with /pol/ . They'd rather cling to what they know than embrace change, or get in rhetorical zingers, and nothing wi…

Counterpoint: in my consulting role, I've directly seen well over a billion dollars in failed AI deployments in enterprise environments. They're good at solving narrow problems, but fall apart in problem spaces exceeding roughly thirty concurrent decision points. Just today I got involved in a client's data migration where the agent (Claude) processed test data instead of the intended data identified in the prompt. It went so far as to rename the test files to match the actual source data files and proceed from there, signalling the all clear as it did. It wasn't caught until that customer, in a workshop said, and I quote "This isn't our fucking data".

Re: DeepMind and OpenAI win gold at ICPC

#163
post #29
post #22

Given that ICPC problems are in general easier than IOI problems. I wouldn't be surprise to see they can get Gold (even perfect scores) in ICPC. Nonetheless, I'm still questioning what's the cost and how long it would take for us to be able to access these models. Still great work, but it's less useful if the cost is actually higher than hiring someone with the same level.

Not sure by what metric you compare the difficulty, but regardless of the hardness of the problem, IIRC, ICPC requires 100% correctness on test cases to score a problem (even failing one means you don't get the score,) but IOI would admit fractional scores (correct me if I am wrong.)

[deleted]

Re: DeepMind and OpenAI win gold at ICPC

#164
post #59

I've contemplated this a bit, and I think I have a bit of an unconventional take: First, this is really impressive. Second, with that out of the way, these models are not playing the same game as the human contestants, in at least two major regards. First, and quite obviously, they have massive amounts of compute power, which is kind of like giving a human team a week instead of five hours. But the models that are co…

Firstly, automobiles are really impressive. Second, with that out the way, these cars are not playing the same game as horses… first, and quite obviously they have massive amounts of horsepower, which is kind of like giving a team of horses… many more horses. But also cars have an absolutely massive fuel capacity. Petrol is such an efficient store of chemical energy compared to hay and cars can store gallons of it. I…

Comparing power with reasoning does not make any sense at all.

Humans have surpassed their own strength since the invention of the lever thousands of years ago. Since then, it has been a matter of finding power sources millions of times greater such as nuclear energy

Re: DeepMind and OpenAI win gold at ICPC

#165

Earlier quoted context omitted.

Firstly, automobiles are really impressive. Second, with that out the way, these cars are not playing the same game as horses… first, and quite obviously they have massive amounts of horsepower, which is kind of like giving a team of horses… many more horses. But also cars have an absolutely massive fuel capacity. Petrol is such an efficient store of chemical energy compared to hay and cars can store gallons of it. I…

Yeah man, and it would be wild to publish an article titled "Ford Mustang and Honda Civic win gold in the 100 meter dash at the Olympics" if what happened was the companies drove their cars 100 meters and tweeted that they did it faster than the Olympians had run. Actually that's too generous, because the humans are given a time limit in ICPC, and there's no clear mapping to say how the LLM's compute should be limite…

> what happened was the companies drove their cars 100 meters and tweeted that they did it faster than the Olympians had run

That would be indeed an interesting race around the time cars were invented. Today that would be silly, since everyone knows what cars are capable of, but back then one can imagine a lot more skepticism.

Just as there is a ton of skepticism today of what LLMs can achieve. A competition like this clearly demonstrates where the tech is, and what is possible.

> there's no clear mapping to say how the LLM's compute should be limited to make a comparison

There is a very clear mapping of course. You give the same wall clock time to the computer you gave to the humans.

Because what it is showing is that the computer can do the same thing a human can under the same conditions. With your analogy here they are showing that there is such a thing as a car and it can travel 100 meters.

Once it is a foregone conclusion that an LLM can solve the ICPC problems and that question has been sufficiently driven home to everyone who cares we can ask further ones. Like “how much faster can it solve the problems compared to the best humans” or “how much energy it consumes while solving them”? It sounds like you went beyond the first question and already asking these follow up questions.

Re: DeepMind and OpenAI win gold at ICPC

#166
post #130

ICPC = The International Collegiate Programming Contest. These are college level programmers, not elite competitive programmers. Apparently Gemini solved one problem (running on who knows what kind of cluster) by burning 30 min of "thinking" time on it, and at a cost that Google have declined to provide. According to one prior competition paricipant, writing in the comments section of this ArsClasica coverage, each y…

The ICPC has plenty of elite competitive programmers. It's an activity that "peaks" in importance around college, and not many keep training a lot after participating. Every year there are multiple "Legendary Grandmasters" in the competition. That's >3000 Elo in Codeforces. I'd estimate it takes a similar level of skill/effort as becoming a Chess Grandmaster. And even those that aren't at that level are very competen…

I didn't realize - didn't mean to disparage anyone.

Re: DeepMind and OpenAI win gold at ICPC

#167

Earlier quoted context omitted.

Deleted, because the "AI" geniuses and power users pointed out that Tao does not have a point. You can get this one to -4 as well, since that seems to be the primary pleasure for "AI" one armed bandit users.

It doesn't say anywhere that Gemini used any of those things at ICPC, or that it used more real-world time than the humans. Also, who cares? It's a self contained non-human system that could solve an ICPC problem it hasn't seen before on its own, which hasn't been achieved before. If there was a savant human contestant with photographic memory who could remember every previous ICPC problem verbatim and can think real…

I think "hasn't seen before" is a bit of an overstatement. Sure, the problem is new in the literal sense that it does exist verbatim elsewhere, but arguably, any competition problem is hardly novel: they are all some permutation of problems that exist and have been solved before: pathfinding, optimization, etc. I don't think anyone is pretending to break new scientific ground in 5 hours.

Re: DeepMind and OpenAI win gold at ICPC

#168

Earlier quoted context omitted.

Yeah man, and it would be wild to publish an article titled "Ford Mustang and Honda Civic win gold in the 100 meter dash at the Olympics" if what happened was the companies drove their cars 100 meters and tweeted that they did it faster than the Olympians had run. Actually that's too generous, because the humans are given a time limit in ICPC, and there's no clear mapping to say how the LLM's compute should be limite…

> what happened was the companies drove their cars 100 meters and tweeted that they did it faster than the Olympians had run That would be indeed an interesting race around the time cars were invented. Today that would be silly, since everyone knows what cars are capable of, but back then one can imagine a lot more skepticism. Just as there is a ton of skepticism today of what LLMs can achieve. A competition like thi…

You're right, they did limit to 5 hours and, I think, 3 models, which seems analogous at least.

Not enough to say they "won gold". Just say what actually happened! The tweets themselves do, but then we have this clickbait headline here on HN somehow that says they "won gold at ICPC".

Re: DeepMind and OpenAI win gold at ICPC

#169
post #59

I've contemplated this a bit, and I think I have a bit of an unconventional take: First, this is really impressive. Second, with that out of the way, these models are not playing the same game as the human contestants, in at least two major regards. First, and quite obviously, they have massive amounts of compute power, which is kind of like giving a human team a week instead of five hours. But the models that are co…

> whereas the teams are allowed to bring a 25-page PDF This is where I see the biggest issue. LLMs are first-and-foremost text compression algorithms . They have a compressed version of a very good chunk of human writing. After being text compression engines, LLMs are really good at interpolating text based on the generalization induced by the lossy compression. What this result really tells us is that, given a reaso…

If we develop a system that can:

- compress (in a relatively recoverable way) the entire domain of human knowledge

- interpolate across the entire domain of human knowledge

- draw connections or conclusions that haven't previously been stated explicitly

- verify or disprove those conclusions or connections

- update its internal model based on that (further expanding the domain it can interpolate within)

Then I think we're cooking with gasoline. I guess the question becomes whether those new conclusions or connections result in a convergent or divergent increase in the number of new conclusions and connections the model can draw (e.g. do we understand better the domains we already know or does updating the model with these new conclusions/connections allow us to expand the scope of knowledge we understand to new domains).

Re: DeepMind and OpenAI win gold at ICPC

#170

Earlier quoted context omitted.

It doesn't matter how many instances were running. All that matters is the wall clock time and the cost. The fact that they don't disclose the cost is a clue that it's probably outrageous today. But costs are coming down fast. And hiring a team of these guys isn't exactly cheap either.

Human teams are limited to three people. So why doesn’t it matter how many instances they used?

This is what the argument is? 10 years ago if you said you could do this with every computer on the planet and every computer scientist focused on trying to create the code to do this I would’ve given you absurd odds against it getting 12 problems right on ICPC. 10 years ago it couldn’t even reliably parse the question statement.
Post reply on HN