Seems intuitively sound; a larger model would have the ability to differentiate among a larger variety of concepts, which translates to a larger vocabulary and greater ability to use expressive tools such as imagery, metaphor etc etc. I could go on, but brevity is virtuous.
None of this has anything to do with the paper, which is concerned with theoretical computer science and constructs artificial "languages" that have a small representation as a(n idealized theoretical) transformer but whose smallest representation in some other formalisms is much larger. In other words, its conception of succinctness is almost diametrically opposite of the way you appear to have understood it. They l…
Re: Transformers Are Inherently Succinct (2025)
#11Had another read - you’re absolutely right, thanks for the kind correction and explanation.