Live data from Hacker News

Three Major Physics Discoveries and Counting

quantamagazine.org

1–10 of 18 posts

Re: Three Major Physics Discoveries and Counting

#4

Interestingly, she mentions using machine learning for analysis of the Higgs Boson data. Does anyone here know more about this?

In my opinion, CERN is one of the few groups out there that can call their data set "big data" without talking out of their ass. (If it can fit in RAM, it ain't Big Data. Multiple TBs fit in RAM) Their detectors produce something on the order of a petabyte per second which is then pared back immensely to become something that's actually storable. Most of the machine learning I've heard of involving CERN is in reducing that data stream and then highlighting "interesting" things for researchers to take a look at.

Re: Three Major Physics Discoveries and Counting

#5

Interestingly, she mentions using machine learning for analysis of the Higgs Boson data. Does anyone here know more about this?

Not exactly an answer to your question, but there is a Kaggle competition currently in progress in this area, sponsored by CERN: https://www.kaggle.com/c/trackml-particle-identification.

Re: Three Major Physics Discoveries and Counting

#6

A intriguing remark near the end: We are also very excited about the near future, because we plan to start using quantum computing to do our data analysis.

It was nice to see this comment, because to me it means that she is always looking to the future & optimistic.

One of the near-term applications of quantum computing is for machine learning. For example, the D-wave chip is for optimising a binary fitness function. We will soon see if it works out.

Re: Three Major Physics Discoveries and Counting

#7
post #4

Interestingly, she mentions using machine learning for analysis of the Higgs Boson data. Does anyone here know more about this?

In my opinion, CERN is one of the few groups out there that can call their data set "big data" without talking out of their ass. (If it can fit in RAM, it ain't Big Data. Multiple TBs fit in RAM) Their detectors produce something on the order of a petabyte per second which is then pared back immensely to become something that's actually storable. Most of the machine learning I've heard of involving CERN is in reducin…

> If it can fit in RAM, it ain't Big Data. Multiple TBs fit in RAM

To be fair, the price for that RAM starts to steepen at (if not before) the 4TB mark, and 1.5TB might have been the limit on frugal main memory as recently as a year and a half ago.

OTOH, SSDs are very fast even at low cost, so I'd argue if it can fit in directly-attached storage whose aggregate bandwidth compares with the RAM's bandwidth, it ain't Big Data, either. Though even a single PB might be big enough, if the CPUs are too slow.

Re: Three Major Physics Discoveries and Counting

#8

A intriguing remark near the end: We are also very excited about the near future, because we plan to start using quantum computing to do our data analysis.

Using subattomic computers to learn about subattomic particles. That will lead to better subattomic computers which will lead to... (A virtuous cycle like Moore's Law?)

Re: Three Major Physics Discoveries and Counting

#9

Interestingly, she mentions using machine learning for analysis of the Higgs Boson data. Does anyone here know more about this?

We (Member of the ATLAS Experiment, 2008-2016) used Neural networks for trigger decisions (to record or ignore a collision) and Boosted Decision Trees were the big thing among many ATLAS physicists back around 2008-9, so that was also used quite a bit. For the experiments at the LHC you can consider the actual analysis of data a Multivariate Hypothesis Testing exercise. The thing is a counting experiment, you have a theory that provides a prediction, simulates its phenomenological effects, the reactions energy deposition in the detector(s), the electronic signal paths, and then we would run "reconstruction" basically turning electronic readouts into particle trajectories, particle energies and particle types. Under the laws of physics, these clues can be combined to measure the rest mass of initial particles that have long decayed (like the Higgs). Now given difficulty of separating the multiple particles leaving energy behind, it is quite difficult to separate the "background" from known physics from the interesting "signal" from a new theory under test. Rather than having to manually do the analysis, Machine Learning is applied at multiple stages. This can be to do particle identification (is it a muon, an electron or something else) or to maximise the binary seperation between two classes such as the signal/background.

In genereal it is wise to know that Machine Learning, Big Data and Cloud computing has been used in particle physics for decades, but with the LHC a world-wide infrastructure has been created to capture all the learnings that the mainstream are only beginning to discover now. For instance a main paradigm in the analysis model is to move calculation to the data, rather than the other way around, due to the large amounts of data. You may call it MapReduce, we call it physics analysis (Map you statistical analysis across decentralised data, reduce the output through distributed merge jobs, plot and publish). Sorry to sound like an old fart, you question is honest and relevant, but it really underlines how easy a story about how Google/Facebook/whatever invented something can rewrite history. Most of the stuff people in the IT/Tech sector are playing around with are inspired by basic or applied science, and applied in a commercial setting. This is exactly how it is supposed to work, but damned if the log analysing marketeers at Google should have all the credit for these developments :)

Now, with my rant over, here are a few references that may be interesting to you:

These were the tools used for physics analysis:

http://tmva.sourceforge.net https://root.cern.ch https://twiki.cern.ch/twiki/bin/view/RooStats/WebHome

And a few articles http://atlas.cern/search/node/Boosted%20Decision%20tree https://cds.cern.ch/search?ln=en&sc=1&p=Machine+Learning&act...

Oh and a bit of gossip. We called Sau Lan Wu the "Dragon lady" (mostly behind her back), because of her awesome energy and tenacity. She really deserves the credit given in the article!

Re: Three Major Physics Discoveries and Counting

#10
post #4

Interestingly, she mentions using machine learning for analysis of the Higgs Boson data. Does anyone here know more about this?

In my opinion, CERN is one of the few groups out there that can call their data set "big data" without talking out of their ass. (If it can fit in RAM, it ain't Big Data. Multiple TBs fit in RAM) Their detectors produce something on the order of a petabyte per second which is then pared back immensely to become something that's actually storable. Most of the machine learning I've heard of involving CERN is in reducin…

Although LHC experiments are toward the top of the heap in terms of data rates, many particle and astronomical experiments produce "big data". One quarter of the DUNE experiment will acquire about 50 EB/year (that's exabyte), outputting about 10 PB/year to tape. LSST will produce data in the few PB/year range.
Post reply on HN