Kolmogorov-Arnold Networks
github.com
Kolmogorov-Arnold Networks
1–10 of 149 posts
Re: Kolmogorov-Arnold Networks
#2I wonder how many more new architectures are going to be found in the next few years
Re: Kolmogorov-Arnold Networks
#3Would this approach (with non-linear learning) still be able to utilize GPUs to speed up training?
Re: Kolmogorov-Arnold Networks
#4Interesting! Would this approach (with non-linear learning) still be able to utilize GPUs to speed up training?
Re: Kolmogorov-Arnold Networks
#5Re: Kolmogorov-Arnold Networks
#6Re: Kolmogorov-Arnold Networks
#7Re: Kolmogorov-Arnold Networks
#8It’d be really cool to see a transformer with the MLP layers swapped for KANs and then compare its scaling properties with vanilla transformers
Given its sparse, Will this be just replacement for MoE.
Re: Kolmogorov-Arnold Networks
#9I can't assess this, but I do worry that overnight some algorithmic advance will enhance LLMs by orders of magnitude and the next big model to get trained is suddenly 10,000x better than GPT-4 and nobody's ready for it.
I would worry if I'd own Nvidia shares.
Re: Kolmogorov-Arnold Networks
#10I can't assess this, but I do worry that overnight some algorithmic advance will enhance LLMs by orders of magnitude and the next big model to get trained is suddenly 10,000x better than GPT-4 and nobody's ready for it.