Live data from Hacker News

Kolmogorov-Arnold Networks

github.com

111–120 of 149 posts

Re: Kolmogorov-Arnold Networks

#111
post #107

Earlier quoted context omitted.

There's a ton actually. Just they tend to go through extra rounds of review (or never make it...) and never make it to HN unless there's special circumstances (this one is MIT and CIT). Unfortunately we've let PR become a very powerful force (it's always been a thing, but seems more influential now). We can fight against this by up voting things like this and if you're a reviewee, not focusing on sota (it's clearly b…

For example, I find Spike Neural Networks to be cool, but until they reach SOTA, how can they displace conventional neural networks?

Compare how much time has been spent studying the two different architectures. Who knows if SNNs can displace other stuff, but I wouldn't rely on SOTA for being the benchmark. Progress has to be made and it isn't made in leaps and bounds. If you find them cool, study them more. Maybe you'll stumble onto something. Maybe you'll find an edge in a niche domain (and maybe you find that that edge can generalize more than you initially thought).

Stop worrying about displacing conventional networks and start worrying about understanding things. We chip away at this together, as a community. There's a lot we need to learn and a lot that needs to be explored. Why tie anyone's hands behind their backs?

Re: Kolmogorov-Arnold Networks

#112
post #107

Earlier quoted context omitted.

For example, I find Spike Neural Networks to be cool, but until they reach SOTA, how can they displace conventional neural networks?

Compare how much time has been spent studying the two different architectures. Who knows if SNNs can displace other stuff, but I wouldn't rely on SOTA for being the benchmark. Progress has to be made and it isn't made in leaps and bounds. If you find them cool, study them more. Maybe you'll stumble onto something. Maybe you'll find an edge in a niche domain (and maybe you find that that edge can generalize more than…

I'm no stranger to having written papers that follow my own curiosity that didn't show any promising results.

However, I wouldn't blame "the community" for not taking my idea and building on it. There needs to be a seed of hope, a taste of future benefits, or else why is it anybody's obligation to care about something subpar?

The introducer of a novel idea needs to beat the incumbent by a large margin. This is just reality, not injustice.

Re: Kolmogorov-Arnold Networks

#114

I quickly skimmed the paper, got inspired to simplify it, and created some Pytorch Layer : https://github.com/GistNoesis/FourierKAN/ The core is really just a few lines. In the paper they use some spline interpolation to represent 1d function that they sum. Their code seemed aimed at smaller sizes. Instead I chose a different representation, aka fourier coefficients that are used to interpolate the functions of indiv…

When I played around with implementing this last night I found using a radial basis function instead of Fourier coefficients (I tried the same, nice and parallel and easy to write) to be more well behaved in training networks of depth greater than 2.

Re: Kolmogorov-Arnold Networks

#115
post #96

Earlier quoted context omitted.

Sorry to hear all this (after writing my other sibling comment). Please don’t lose faith in the review process. It is still useful. Until the AGI can be better reviewers, which is hopefully not too far in the future.

For me to regain faith in the review process I need to actually see some semblance of the review process working. So far, instead, I've seen: - Banning social media posting so that only big tech and collusion positing can happen to "protect the little guy" - Undoing the ban to lots of complaints - Instituting a no LLM policy with no teeth and no method to actually verify - Instituting a high school track to get those…

I feel you. Here are some thoughts from the other side of the fence:

Social media banning aims to preserve anonymity when the reviews are blind. It is hard to convincingly keep anonymity for many submissions, but an effort to keep it is still worthwhile and typically helps the less privileged to get a fair shot at a decent review, avoiding the social media popularity contest.

The policies for LLM usage differ between conferences. The only possibly valid concern with use of AI is the disclosure of non public info to an outside LLM company that may happen to publish or be retrained on that data (however unlikely this is in practice) before the paper becomes public; for example, someone could withdraw their publication and it no longer sees the day of light on the openreview website. (I personally disagree with this concern.) As far as I know there is no real limitation to using self hosted AI as long as the reviewer takes full credit for the final product and there is no limitation in using non public AI to improve the review clarity without dumping the full paper text. A fraction of authors would appreciate better referee reports, so at a minimum, the use of AI can bridge the language gap. I wouldn’t mind the conferences instituting an automatic AI processing to help the reviewers reduce ambiguity and avoid trivialities.

The high school track has been ridiculed, as expected. I think it is a great idea and doesn’t only apply to rich kids. There exist excellent specialized schools in NYC and other places in the US that might find ways to get resources for underprivileged ambitious high schoolers. It is possible that in the future a variant of such a track will incentivize some industry to donate compute resources to high school programs and it may start early and powerful local communities. I learned a lot in what would be middle school in the US by interacting with self motivated children at a ad hoc computer club and kept the same level of osmotic learning in the computer lab at college. The current state of AI is not super deep in terms of background knowledge, mostly super broad, and some specialized high schools already cover calculus and linear algebra, and certainly many high schools nowadays provide sufficient background in programming and elementary data analysis.

My personal reward hacking is that the conferences provide a decent way to focus the review to the top hundred or couple hundred plausible abstracts and even when the eventual choice is wrong I get a much better reward to noise ratio than from social media and the pure attacks on the arxiv (although LLMs help here as well). I always find it refreshing to see the novel ideas when they are in a raw form before they have been polished and before everyone can easily judge their worth. Too many of them get unnecessary negative views, which is why the system integrates multiple reviewers and area chairs that can make corrective decisions. It is important to avoid too much noise even at the risk of missing a couple great ones, and yet it always hurts when people drop greatness because of misunderstandings or poor chair choices. No system is perfect, but scaling these conferences from a couple hundred people a year up to about a dozen years ago to approaching hundred thousand a year has worked reasonably well.

Re: Kolmogorov-Arnold Networks

#116
post #15

Earlier quoted context omitted.

Actually, that would be fantastic for NVIDIA shares; 1. A new architecture would make all/most of these upcoming Transformer accelerators obsolete => back to GPUs. 2. Higher performance LLMs on GPUs => we can speed up LLMs with 1T+ parameters. So, LLMs become more useful, so more of GPUs would be purchased.

1. A new architecture would make all/most of these upcoming Transformer accelerators obsolete => back to GPUs. There's no guarantee that that is what would happen. The right (or wrong, depending on your POV) algorithmic breakthrough might make GPU's obsolete for AI, by making CPU's (or analog computing units, or DSP's, or "other") the preferred platform to run AI.

Assuming there is a development that makes GPUs obsolete, I think it's safe to assume that what will replace them at scale will still take the form dedicated AI card/rack

1. Tight integration necessary for fundamental compute constraints like memory latency.

2. Economies of scale

3. Opportunity cost to AI orgs. Meta, OpenAI etc want 50k h100s to arrive in shipping container and plug in so they can focus on their value-add.

Everyone will have to readjust to this paradigm. Even if next get AI runs better on CPU, Intel won't suddenly be signing contracts to sell 1,000,000 xeons and 1,000,000 motherboards etc

Also, Nvidia have 25bn cash in hand and almost 10 billion yearly r&d spend. They've been an AI-first company for over a decade now, they're more prepared to pivot than anyone else

Edit: nearly forgot - Nvidia can issue 5% new stocks and raise 100B like it's nothing.

Re: Kolmogorov-Arnold Networks

#117
post #94

Earlier quoted context omitted.

I agree with almost all you said except that Twitter is better than top conferences, and I take a contrarian view that reviewers slow down AGI with requests for additional experiments. Without going into specifics, which you can probably guess based on your background, too many ideas that work well, even optimally, at small scale fail horribly at large scale. Other ideas that work at super specialized settings don’t…

> too many ideas that work well, even optimally, at small scale fail horribly at large scale. Not that I disagree, but I don't think that's a reason to not publish. There's another way to rephrase what you've said many ideas that work well at small scales do not trivially work at large scales But this is true for many works, even transformers. You don't just scale by turning up model parameters and data. You can, but…

I think information gain will be easy to measure in principle with an AI in the near future: if the work is correct, how unexpected is it. Anything trivially predictable based on published literature, including exact reproduction disguised as novel is not worthy of too much attention. Anything that has a change of changing the model of the world is important. It can seem minor even trivial to some nasty reviewer, but if the effect is real and not demonstrated before then it deserves attention. Until then, we deal with imperfect humans.

Regarding large multimodal data, I don’t know what people you refer to, so I can’t comment further. The current math is useful but very limited when it comes to understanding the densities in such data; vectors are always orthogonal at high dim and densities are always sampled very poorly. The type of understanding of data that would help progress in drug and material design, say, is very different from the type of data that can help a chatbot code. Obviously the future AI should understand it all, but it may take interdisciplinary collaborations that best start at an early age and don’t fit the current academic system very well unfortunately.

Re: Kolmogorov-Arnold Networks

#118
post #5

It’d be really cool to see a transformer with the MLP layers swapped for KANs and then compare its scaling properties with vanilla transformers

This is the first thought came to my mind too. Given its sparse, Will this be just replacement for MoE.

MoE is mostly used to enable load balancing since it makes it possible to put experts on different GPUs. This isn't so easy to do with a monolithic, but sparse layer.

Re: Kolmogorov-Arnold Networks

#119
post #112

Earlier quoted context omitted.

Compare how much time has been spent studying the two different architectures. Who knows if SNNs can displace other stuff, but I wouldn't rely on SOTA for being the benchmark. Progress has to be made and it isn't made in leaps and bounds. If you find them cool, study them more. Maybe you'll stumble onto something. Maybe you'll find an edge in a niche domain (and maybe you find that that edge can generalize more than…

I'm no stranger to having written papers that follow my own curiosity that didn't show any promising results. However, I wouldn't blame "the community" for not taking my idea and building on it. There needs to be a seed of hope, a taste of future benefits, or else why is it anybody's obligation to care about something subpar? The introducer of a novel idea needs to beat the incumbent by a large margin. This is just r…

The incumbent approaches usually benefit from a ton of research that might or might not be transferrable to the newcomer.

Even if many optimizations also apply to the new approaches, taking advantage of them takes a lot of work. For example, I have not yet implemented KV caches for my nanoGPTs that I'm fooling around with.

Re: Kolmogorov-Arnold Networks

#120

Earlier quoted context omitted.

How GPU-friendly is this class of models?

Very unfriendly. The symbolic library (type of activations) requires a branching at the very core of the kernel. GPU will need to serialized on these operations warp-wise. To optimize, you might want to do a scan operation beforehand and dispatch to activation funcs in a warp specialized way, this, however, makes the global memory read/write non-coalesced. You then may sort the input based on type of activations and…

couldn't you implement these as a texture lookup, where x is the input and the various functions are stacked in y? That should be quite fast on gpus.
Post reply on HN