There are so many great mathematical PDFs available for free on the Internet, written by academics/educators/engineers. A problem is that there is a huge amount of overlap. I wonder if an AI model could be developed that would do a really good job of synthesizing an overlapping collection into a coherent single PDF without duplication.
Learning Theory from First Principles [pdf]
11–20 of 66 posts
Re: Learning Theory from First Principles [pdf]
#12I can’t wait until I tell GPT-5 “I have this idea I want to try, read this book and tell me there’s anything relevant there to make it work better”.
Re: Learning Theory from First Principles [pdf]
#13I can’t wait until I tell GPT-5 “I have this idea I want to try, read this book and tell me there’s anything relevant there to make it work better”.
1. It is very difficult for me to tell you about my context as a user within low dimension variables.
2. I do not understand my situation in the universe to be able to tell AI.
3. I dont have a vocabulary with AI. Internet i feel aced this with shared HTTP protocol to consistently share agreed upon state. For ex within Uber I am a very narrow request response universe with.. POST phone, car, gps(a,b,c,d), now, payment.
But as a student wanting to learn algorithms how do I pass that I'm $age $internet-type from $place and prefer graphical explanations of algorithms, have tried but gotten scared of that thick book and these $milestones-cs50, know $python upto $proficiency(which again is a fractal variable with research papers on how to define for learning).
Similarly how do I help you understand what stage my startup idea is beyond low traction, but want to know have $networks/(VC, devs, sales) APIs, have $these successful partnerships with such evidence $attendance, $sales. Who should I speak to? Could you pls write the needful in mails and engage in partnership with other bots under $budget.
Even in the real world this vocabulary is in smaller pockets as our contexts are too different.
4. Learning assumes knowledge exists as a global forever variable in a wider than we understand universe. $meteor being a non maskable interrupt to the power supply at unicorn temperatures in a decade. Similarly one time trends in disposable $companies that $ecosystem uses to learn. I'm in a desert village with with absent electricity might mean those machines never reach me and perhaps most people don't have a basic phone in the world to be able to share state. Their local power mafia politics and absent governance might mean the pdf AI recommends i read might or might not help.
I don't know how this will evolve but to think of the possibilities has been so interesting. It's like computers can talk to us easily and they're such smart babies on day 1 and "folks we aren't able to put right, enough, cheap data in" is perhaps the real bottleneck to how much usefulness we are being able to uncover.
Re: Learning Theory from First Principles [pdf]
#14I can’t wait until I tell GPT-5 “I have this idea I want to try, read this book and tell me there’s anything relevant there to make it work better”.
Ypur idea would become even better so if you read the book yourself ... You might even learn something new along the way.
Re: Learning Theory from First Principles [pdf]
#15A mention of no free lunch theorem should come with a disclaimer that the theorem is not relevant in practice. An assumption that your data originates from the real world, is sufficient that the no free lunch theorem is not a hindrance.
This book doesn't discuss this at all. Maybe mention that "all distributions" means a generalization to higher dimensional spaces of discontinuous functions (including the tiny subset of continuous functions) of something similar to all possible bit sequences generated by tossing a coin. So basically if you data is generated from an even random distribution of "all possibilities", you cannot learn to predict the outcome of the next coin tosses, or similar.
Re: Learning Theory from First Principles [pdf]
#16Interesting! I’ll have to look over it when I have more time. From a quick glance, it looks like it covers much of the same material as this text [1]. I wonder how they compare. [1]: https://www.cambridge.org/core/books/understanding-machine-l...
Re: Learning Theory from First Principles [pdf]
#17> 2.5 No free lunch theorem > > Although it may be tempting to define the optimal learning algorithm that works optimally for all distributions, this is impossible. In other words, learning is only possible with assumptions. A mention of no free lunch theorem should come with a disclaimer that the theorem is not relevant in practice. An assumption that your data originates from the real world, is sufficient that the n…
For example, many people naively think that static program analysis is unfeasible due to the halting problem, Rice's theorem, etc.
Re: Learning Theory from First Principles [pdf]
#18Have they figured out what causes double descent yet?
Re: Learning Theory from First Principles [pdf]
#19Have they figured out what causes double descent yet?
Here a "feature" might be seen as an abstract, very, very high dimensional vector space. The team is pretty deep in investigating the idea of superposition, where individual neurons encode for multiple concepts. They experiment with a toy model and toy data set where the latent features are represented explicitly and then compressed into a small set of data dimensions. This forces superposition. Then they show how that superposition looks under varying sizes of training data.
It's obviously a toy model, but it's a compelling idea. At least for any model which might suffer from superposition.
https://transformer-circuits.pub/2023/toy-double-descent/ind...
Re: Learning Theory from First Principles [pdf]
#20I can’t wait until I tell GPT-5 “I have this idea I want to try, read this book and tell me there’s anything relevant there to make it work better”.
I've never heard of an LLM making up a new idea. Shouldn't this only work if your thing has already been tried before?