> [them] Output tokens: FREE (too cheap to meter).
I'm very confused by this.
21–30 of 513 posts
> [them] Output tokens: FREE (too cheap to meter).
I'm very confused by this.
edit: looks like a framer export where there is a text stroke being applied :|
It could be used for coding if you gave it an AST. If you work at TypeSafe please try this. Side note: This is probably how LLMs would perform with better encoders and next-latent prediction, so eventually those will beat this architecture out. Still amazing though.
I'd love to do research on this when I have the time.
> Extraordinary claims require extraordinary evidence so see below for the receipts. Yes, that’s the kind of attitude I want to see in these model releases
Earlier quoted context omitted.
Direct link to the Doom video tweet: https://x.com/completeskeptic/status/2099925687465570372
The doom video is also in the article itself (headline: "Doom"). I suppose this is the same video as the one from the parent comment, but I don't know for sure - I don't have a twitter account and the above link doesn't work for me.
Does this imply it's a very small model? I couldn't find anything about the model itself.
um what is going on with the outfit changes in the launch video... https://x.com/CompleteSkeptic/status/2099925682726002904
> Extraordinary claims require extraordinary evidence so see below for the receipts. Yes, that’s the kind of attitude I want to see in these model releases
But the evidence is not there...
This sounds good but so far all claims just sound like marketing terms. I'd love to see real proof. e.g. "RLCD" and "parallel sampling" have nothing to back it up. also "70-500ms vs 3-329 seconds" are apples-to-oranges unless the LLM baseline is doing comparable work (e.g., long chain-of-thought). If Jev is skipping generation entirely for a narrow structured task, of course it's faster. Nonetheless i want this to be…
It's totally reasonable to compare against LLMs doing chain of thought if it gets comparable performance.