Live data from Hacker News

GPT-3 is no longer the only game in town

lastweekin.ai

191–200 of 217 posts

Re: GPT-3 is no longer the only game in town

#191
post #93

The future is not as dark as it seems because of the rat race of megacorps. You can use reduced versions of language models with extremely good results. I was involved in training the first-ever GPT2 for Bengali language, but with 117 million parameters. It took a month's effort (training + writing code + setup) and about $6k in TPU cost, but Google Cloud covered it. Anyway, it is surprisingly good. We fine-tuned the…

That is a fantastic result - nagging question - these work best on predictable things. How much of Bengali poetry is predictable?

> these work best on predictable things

Umm, not really.

You are talking about single and multi-label classification tasks, maybe?

Bengali poetry is just like poetry in any other rich languages like English, French, etc.

What these models do, from a high level, is that they learn the distribution of the data. In this case, they learn the style of the poets.

Writing poetry in specific styles has been done long before Transformer architectures came.

See-

- https://www.tensorflow.org/text/tutorials/text_generation

- https://machinelearningmastery.com/text-generation-lstm-recu...

The goal of poetry generation is to generate something that is unique, is in the poetic style, and is coherrent, grammaticallly correct, ideally indistinguishable in the eyes of a human.

Re: GPT-3 is no longer the only game in town

#192
post #93

The future is not as dark as it seems because of the rat race of megacorps. You can use reduced versions of language models with extremely good results. I was involved in training the first-ever GPT2 for Bengali language, but with 117 million parameters. It took a month's effort (training + writing code + setup) and about $6k in TPU cost, but Google Cloud covered it. Anyway, it is surprisingly good. We fine-tuned the…

> about $6k in TPU cost, but Google Cloud covered it. I'm glad this all worked out for you. This is unrelated, but I just want to say that I hate how many people Google managed to convert to TPU with their research program and that their managed TPU/GPU offerings are absolutely horrible and infuriating to work with unless you somehow get on their radar.

I have never used TRC, but I heard from many that if you reach out to them, they are really helpful.

I have not converted to TPU, because it is literally offered by only one company. It will be the height of "vendor lock-in".

But, I must say that JAX on TPU is faster than anything and everything that I have ever seen.

Re: GPT-3 is no longer the only game in town

#193

I never got why GPT-3 was so closed off, like you needed permission to use it. If it’s so good then why not just make it available?

> If it’s so good then why not just make it available?

OpenAI is pivoting to corporate evil, and to do that properly they need proprietary assets to rent out.

Re: GPT-3 is no longer the only game in town

#194

Earlier quoted context omitted.

“I would challenge you to name any (non-simple) problem where traditional AI methods are still state of the art.” Lossless file compression. As far as I know none of the algorithms in widespread use are neural-based, despite the fact that compression is clearly a rich statistical modeling problem, at least on par with GPT-3-style language understanding in difficulty. There are published attempts to solve the problem…

I'm far from an expert in this subject but doesn't this ranking of large text compression algorithms with NNCP coming first suggest that neural-nets are pretty great at compression? http://mattmahoney.net/dc/text.html https://bellard.org/nncp/ I don't see examples of high performing symbolic AI based compression algorithms anywhere, but again I am very ignorant, do you have examples?

The ranking criteria of this list make it very unrepresentative of compressors used in the real world. The benchmark they’re using for example is the sum of the compressed file plus the compressor binary: this penalizes memorization of the evaluation text in the compressor binary itself. But in the real world, you would have no concerns at all that your compressor is “cheating” by working too well only for your particular data — having useful priors that model real-world data for more compact representations is the whole point. Many of these algorithms are also impractical due to speed or memory use. Ask yourself: How many of the top-10 algorithms do you have installed right now, or even recognize? The winners aren’t dominating outside the arena of this list.

I’m also not an expert in symbolic AI — my comment above is more about neural vs. pre-neural NLP methods, rather than symbolic AI, which I admit drifts a bit from the parent. A compressor replacing word tokens with dictionary indices is definitely symbolic but it’s not especially “AI”.

Re: GPT-3 is no longer the only game in town

#195

Earlier quoted context omitted.

PHP wasn't discredited by the incumbents. It was discredited by its creator. "I'm not a real programmer. I throw together things until it works then I move on. The real programmers will say Yeah it works but you're leaking memory everywhere. Perhaps we should fix that. I'll just restart Apache every 10 requests." -Rasmus Lerdorf "I was really, really bad at writing parsers. I still am really bad at writing parsers."…

To most programmers that doesn't discredit PHP at all. He cares about a working product, much like 90% of programmers, who don't have the privilige to worry about theory. They just need an ecommerce, or blog or whatever, running asap. To use a pg's analogy, they are there to paint not to worry about painting chemistry. The incumbents do discredit PHP though. For instance, facebook was built on PHP, and still runs on…

PHP is also discredited by its apologists.

If you're just there to paint, but painting with mashed potatoes, you SHOULD have been more worried about your paint chemistry.

"Using these toolkits is like trying to make a bookshelf out of mashed potatoes." -JWZ

Re: GPT-3 is no longer the only game in town

#196
post #191

Earlier quoted context omitted.

That is a fantastic result - nagging question - these work best on predictable things. How much of Bengali poetry is predictable?

> these work best on predictable things Umm, not really. You are talking about single and multi-label classification tasks, maybe? Bengali poetry is just like poetry in any other rich languages like English, French, etc. What these models do, from a high level, is that they learn the distribution of the data. In this case, they learn the style of the poets. Writing poetry in specific styles has been done long before…

Ah, I see the distinction now, thanks.

It's an interesting subject, I think, since poetry has so many forms it can take and you need the output to capture the idiosyncratic aesthetics and "inner world" of a piece of verse.

I've actually tried using GPT-3 to generate poetry and the naive approach of just sending prompts and text snippets through the API had wildly varying results. Some pretty good, some that were basically word salad and nonsense.

But I also run a poetry journal on the side, so maybe that skews my understanding of it!

Re: GPT-3 is no longer the only game in town

#197
post #99

Earlier quoted context omitted.

Not at all. Brevity (or verbosity) is largely orthogonal to level of entropy or redundancy. In principle it ought to be possible to code at a higher level of abstraction while still using understandable names and control flow constructs.

I’m really not sure about that. By definition, high entropy means high information density (information per character). So with the same amount of information you would have less characters.

Think in terms of information density per symbol on the abstract syntax tree, not per character in the source code.

Re: GPT-3 is no longer the only game in town

#199
post #114

Earlier quoted context omitted.

Are you confusing libel with something else? Can you extrapolate what you mean here? Are you saying that they will be liable for libel (!) if they publish a negative summary of a product?

If they mischaracterize a positive review into a negative summary based on factual mistakes they know the system makes at a high rate, I would think they would be liable for libel right?

Maybe under UK "judgements widely banned from enforcement" libel laws but it would be basically impossible for ML to fall afoul of it in the US. It could not even be knowingly false. Reckless disregard for the truth would also be hard to argue as it is meant to be a best effort in accuracy.
Post reply on HN