Live data from Hacker News

Jeff Dean responds to EDA industry about AlphaChip

twitter.com

211–220 of 221 posts

Re: Jeff Dean responds to EDA industry about AlphaChip

#211
post #62

Earlier quoted context omitted.

> From where I'm sitting it looks like, "Google spent a fortune on deep learning, and got a small but real win. People who don't like Google failed to follow Google's recipe and got a large and easily replicated loss." From where I'm sitting it looks like Google cooked the books maximally, barely beat humans let alone state of the art algorithms, published a crappy article in Nature because it would never have passed…

> published a crappy article in Nature because it would never have passed editorial muster at something like DAC or an IEEE journal and now have to browbeat other people who are calling them out on it. I don't think it's easier to get into DAC / an IEEE journal than Nature. Their human baseline was the TPU physical design team, with access to the best available tools: rdcu.be/cmedX and this is still the baseline to b…

Nature papers get retracted every year. I have not heard of DAC papers retracted.

If the Nature paper made it clear that RL is not seriously expected to work on non-TPU chips, it would have have probably been rejected. If RL works on many other chips, then evidence should be easy to publish.

Re: Jeff Dean responds to EDA industry about AlphaChip

#212

Earlier quoted context omitted.

> direct comparisons in Cheng That's the ISPD paper referenced many times in this whole thread. > Stronger Baselines Re: "Stronger baselines", the paper "That Chip Has Sailed" says "We provided the committee with one-line scripts that generated significantly better RL results than those reported in Markov et al., outperforming their “stronger” simulated annealing baseline." What is your take on this claim? As for 're…

The point is that the Cheng et al results and paper were shown to Google and apparently okayed by Google points of contact. After this, complaining that Cheng et al didn't ask someone outside Google makes little sense. These far fetched excuses and emotional wording by Jeff Dean leave a big cloud over the Nature work. If he is confident everything is fine, he would not bother. To clarify "you'd expect" - if Jeff Dean…

Could you please point out the specific lines you are dissatisfied with? Is it something an additional publication cannot resolve?

Additionally, in case you forgot to answer, what is your wish for the future of this line of research? Do you hope to see it improve the EDA status quo, or would you prefer the work to stop entirely? If it is the latter, I would have no intention of continuing this conversation.

Re: Jeff Dean responds to EDA industry about AlphaChip

#213

Earlier quoted context omitted.

By what measure are TPUs “successful”? Where is your data coming from?

They're the only non-nvidia accelerator used to train state of the art large language models at scale?

AWS Trainium does as well: https://aws.amazon.com/ai/machine-learning/trainium/

Re: Jeff Dean responds to EDA industry about AlphaChip

#214

Earlier quoted context omitted.

All these papers doing "research" on how to better prompt ChatGPT would be unpublishable then, given that API access to older models gets retired, so the findings of these papers can no longer be reproduced. (I agree with you in principle; my example above is meant to show that standards for things such as reproducibility aren't easily defined. There are so many factors to consider.)

Well since you put "research" in quotes, I think you also agree that this type of work does not really belong in a quality journal with a high impact factor ;)

This, their training data doesn't even seem to be open either. So it's literally impossible to replicate their model. This makes me highly skeptical.

Re: Jeff Dean responds to EDA industry about AlphaChip

#215

Earlier quoted context omitted.

the link is for a wrongful termination lawsuit, related to the fraud but not a case for the fraud itself. settled may 2024

"Settled" does not mean "Dean did nothing wrong". It means "Google paid the plaintiffs a lot of money so they'd stop saying publicly that Dean did something wrong", which is very different.

[deleted]

Re: Jeff Dean responds to EDA industry about AlphaChip

#216

Earlier quoted context omitted.

"These major methodological differences unfortunately invalidate Cheng et al.’s comparisons with and conclusions about our method. If Cheng et al. had reached out to the corresponding authors of the Nature paper[8], we would have gladly helped them to correct these issues prior to publication[9]. [8] Prior to publication of Cheng et al., our last correspondence with any of its authors was in August of 2022 when we re…

That is misleading. The first two authors left Google in August 2022 under unclear circumstances. The code and data were owned by Google, that's probably why Kahng continued discussibg code and data with his Google contacts. He received clear answers from several Google employees, so if they were at fault, Google should apologize rather than blame Cheng and Kahng.

"Prior to publication of Cheng et al., our last correspondence with any of its authors was in August of 2022 when we reached out to share our new contact information."

You don't stop being the corresponding authors of a paper when you change companies, and whatever "unclear circumstances" you imagine took place when they left, they were also re-hired later, which a company would only do if they were in good standing.

In any case, those "Google contacts" also expressed concerns with how Cheng et al. were doing their study, which they ignored:

3.4 Cheng et al.’s Incorrect Claim of Validation by Google Engineers

Cheng et al. claimed that Google engineers confirmed its technical correctness, but this is untrue. Google engineers (who were not corresponding authors of the Nature paper) merely confirmed that they were able to train from scratch (i.e. no pre-training) on a single test case from the quick start guide in our open-source repository. The quick start guide is of course not a description of how to fully replicate the methodology described in our Nature paper, and is only intended as a first step to confirm that the needed software is installed, that the code has compiled, and that it can successfully run on a single simple test case (Ariane).

In fact, these Google engineers share our concerns and provided constructive feedback, which was not addressed. For example, prior to publication of Cheng et al., through written communication and in several meetings, they raised concerns about the study, including the use of drastically less compute, and failing to tune proxy cost weights to account for a drastically different technology node size.

The Acknowledgements section of Cheng et al. also lists the Nature corresponding authors and implies that they were consulted or even involved, but this is not the case. In fact, the corresponding authors only became aware of this paper after its publication.

Re: Jeff Dean responds to EDA industry about AlphaChip

#217
post #98

Earlier quoted context omitted.

You're saying that if the other methods were given the equivalent amount of compute they might be able to perform as well as AlphaChip? Or at least that the comparison would be fairer? Are the other methods scalable in that way?

The Google internal paper by Chatterjee and the Cheng et al paper from UCSD made such comparisons with Simulated Annealing. The annealer in the Nature paper was easy to improve. When given the same time budget, the improved annealer produced better solutions than AlphaChip. When you give both more time, SA remains ahead. Just read published papers.

The UCSD paper didn't run the Nature method correctly, so I don't see how you can draw this conclusion.

From Jeff's tweet:

"In particular the authors did no pre-training (despite pre-training being mentioned 37 times in our Nature article), robbing our learning-based method of its ability to learn from other chip designs, then used 20X less compute and did not train to convergence, preventing our method from fully learning even on the chip design being placed."

As for Chatterjee's paper, "We provided the committee with one-line scripts that generated significantly better RL results than those reported in Markov et al., outperforming their “stronger” simulated annealing baseline. We still do not know how Markov and his collaborators produced the numbers in their paper."

Re: Jeff Dean responds to EDA industry about AlphaChip

#218

Earlier quoted context omitted.

His original complaint being dismissed matters because it suggests that he was fishing around for a complaint that was valid, and that perhaps his primary motivation was to get money out of Google. Legal nitpick - you can get away with alleging pretty much whatever you want in a legal complaint. You can't even be sued for defamation if it turns out later you were lying. Jeff Dean isn't saying that Cheng et al. should…

No, it doesn't suggest that. Complaints are often dismissed on technicalities or because they are written poorly. Google claimed their new algorithm as a breakthrough. If this were the so, the algorithm would have helped design chips in many different cases. Now, the defense is that it only works for some inputs, and those inputs cannot be shared. This is not a serious defense and looks like a coverup.

Three generations of TPU, Axion (ARM-based CPU), various other chips at AlphaBet, MediaTek's usage...

Re: Jeff Dean responds to EDA industry about AlphaChip

#219
post #200

Earlier quoted context omitted.

When Google published the Nature article, Nature included a rosy intro article by a leading expert in chip design. His name was Andrew Kahng, and he apparently liked Google at the time. But when he dug into Google code (released way after publication), he retracted his intro and co-authored the Cheng et al article. You see how your theory breaks down here.

You are adding information that I did not previously possess. And neither the article nor the previous poster offered. As I said, the truth will out. I was just unhappy with the case previously being offered.

As Andrew Kahng was one of the co-authors of Cheng et al., all of the issues with his reproduction still matter here. The Nature paper went through an investigation and second round of peer review.

AlphaChip is used to make real chips in production. Google publicly announced its use in multiple generations of TPUs and Axion CPUs, and MediaTek said they've built on it as well.

Re: Jeff Dean responds to EDA industry about AlphaChip

#220

Earlier quoted context omitted.

It is open: https://github.com/google-research/circuit_training

As far as I understand it, only kind of? It's open source, but in their paper they did a tonne of pre-training and whilst they've released a small pre-training checkpoint they haven't released the results of the pre-training they've done for their paper. So anyone reproducing this will innevitably be accused of failing to pretrain the model correctly?

I think the pre-trained checkpoint uses the same 20 TPU blocks as the original paper, but it probably isn't the exact-same checkpoint, as the paper itself is from 2020/2021.
Post reply on HN