This is really hard to judge because by the looks of it, finance papers mostly consist of gobbledygook and extensive filler to begin with.
Economics is the attempt to take sociology and add numbers to make it look like a hard science. The fintechbros then seem to think because they can make numbers go up that this proof it's a hard science.
I don't know how you get here from “predict the next word”
31–40 of 275 posts
Re: I don't know how you get here from “predict the next word”
#32It’s called emergent behavior. We understand how an llm works, but do not have even a theory about how the behavior emerges from among the math. We understand ants pretty well, but how exactly does anthill behavior come from ant behavior? It’s a tricky problem in system engineering where predicting emergent behavior (such as emergencies) would be lovely.
> but do not have even a theory about how the behavior emerges We fully do. There is a significant quality difference between English language output and other languages which lends a huge hint as to what is actually happening behind the scenes. > but how exactly does anthill behavior come from ant behavior? You can't smell what ants can. If you did I'm sure it would be evident.
?
Re: I don't know how you get here from “predict the next word”
#33That is my take too, I was surprised to see how many people object to their works being trained on. It's how you can leave your mark, opening access for AI, and in the last 25 years opening to people (no restrictions on access, being indexed in Google).
Re: I don't know how you get here from “predict the next word”
#34> the kind of analysis the program is able to do is past the point where technology looks like magic. I don’t know how you get here from “predict the next word.” You're implicitly assuming that what you asked the LLM to do is unrepresented in the training data. That assumption is usually faulty - very few of the ideas and concepts we come up with in our everyday lives are truly new. All that being said, the refine.in…
This is just as stuck in a moment in time as "they only do next word prediction" What does this even mean anymore? Are we supposed to believe that a review of this paper that wasn't written when that model (It's putatively not an "LLM", but IDK enough about it to be pushy there) was trained? Does that even make sense? We're not in the regime of regurgitating training data (if we really ever were). We need to let go of these frames which were barely true when they took hold. Some new shit is afoot.
Re: I don't know how you get here from “predict the next word”
#35Earlier quoted context omitted.
> but do not have even a theory about how the behavior emerges We fully do. There is a significant quality difference between English language output and other languages which lends a huge hint as to what is actually happening behind the scenes. > but how exactly does anthill behavior come from ant behavior? You can't smell what ants can. If you did I'm sure it would be evident.
> There is a significant quality difference between English language output and other languages ?
Re: I don't know how you get here from “predict the next word”
#36The "predict the next word" to a current llm is at the same level as a "transistor" (or gate) is to a modern cpu. I don't understand llms enough to expand on that comparison, but I can see how having layers above that feed the layers below to "predict the next word" and use the output to modify the input leading to what we see today. It is turtles all the way down.
The next-word bit may be slightly higher than an individual transistor, possibly functional units.
Re: I don't know how you get here from “predict the next word”
#37I have come to think “predict the next token” is not a useful way to explain how LLMs work to people unfamiliar with LLM training and internals. It’s technically correct, but at this point saying that and not talking about things like RLVR training and mechanistic interpretability is about as useful as framing talking with a person as “engaging with a human brain generating tokens” and ignoring psychology. At least A…
I prefer to use the term "spicy autocomplete" myself.
Re: I don't know how you get here from “predict the next word”
#38Earlier quoted context omitted.
> There is a significant quality difference between English language output and other languages ?
They're saying LLMs do better when outputting English than other languages, an assertion I'm not really able to test but have heard elsewhere.
Re: I don't know how you get here from “predict the next word”
#39> Nothing you write will matter if it is not quickly adopted to the training dataset. That is my take too, I was surprised to see how many people object to their works being trained on. It's how you can leave your mark, opening access for AI, and in the last 25 years opening to people (no restrictions on access, being indexed in Google).
Your surprise to people’s objections makes sense if you can’t count.
Re: I don't know how you get here from “predict the next word”
#40Earlier quoted context omitted.
> but do not have even a theory about how the behavior emerges We fully do. There is a significant quality difference between English language output and other languages which lends a huge hint as to what is actually happening behind the scenes. > but how exactly does anthill behavior come from ant behavior? You can't smell what ants can. If you did I'm sure it would be evident.
Two very big revelations here that I would love to know more about: 1. Can you reveal "what's actually happening behind the scenes" beyond the hint you gave? I can't figure it out. 2. Can you explain how an ants sense of smell leads to anthills?
Ant 0: doesn’t seem to be dangerous here. I’ll drop a scent.
Ant 1: oh cool, a safe place. And I didn’t die either. I’ll reinforce that.
Ant 142,857,098,277: cool anthill.