Live data from Hacker News

Alpaca: A strong open-source instruction-following model

crfm.stanford.edu

131–140 of 313 posts

Re: Alpaca: A strong open-source instruction-following model

#131
post #121

I've played a lot with davinci 3 ($25 of credits worth) and it can do some impressive rhyming and interpretation of concepts as emoji sequences. From the 3 times I've interacted with this fine tuned llama 7B it is clear it cannot do that. I've also run the "vanilla" 7B, 13B, and 30B on my home computer with llama.cpp modified for interactive "chat" mode with various pre-prompt and these can't do it either. I have no…

7B parameters is next to nothing when compared to gpt3. If 7B works as well as it does here, A fine tuned 65B model could very easily achieve chatGPT level performance.

I thought ChatGPT is only 20B parameters to begin with?

(Source https://www.forbes.com/sites/forbestechcouncil/2023/02/17/is...)

Re: Alpaca: A strong open-source instruction-following model

#133
post #28

This is why I think we're seeing a Stable Diffusion moment for LLMs: https://simonwillison.net/2023/Mar/11/llama/ Look at the timeline: 24th February 2023: LLaMA is announced, starts being shared with academic partners: https://research.facebook.com/publications/llama-open-and-ef... 2nd March: Someone posts a PR with a BitTorrent link to the models: https://github.com/facebookresearch/llama/pull/73 10th March: First…

Question: what percentage of the hype and momentum for this is so people can run sex chatbots on their local machine?

Re: Alpaca: A strong open-source instruction-following model

#134
post #124

Earlier quoted context omitted.

"Accomodate" is the word to scrutinize here. Yes, it will cost a lot to outright buy physical HPC infrastructure to train and infer a series of large models deployed for customers all over the globe. No, it won't cost nearly as much to rent cloud infra to train a similarly-sized model. No, you won't be able to train a large model on a single multi-GPU node, you will need a cluster containing a respectable power of tw…

You are missing the point. Extremely large LLMs don't train the same way as your BERT_Large x8 variety of LLMs. Your whole training procedure is different. Also Microsoft spent so much initially because their Azure Cloud was unable to cope with it electrically and they had to rewire a datacenter for it. So it's not even a question of just renting 1000 GPUs. Do you have actual experience training GPT-3+ sized models?

If you are interested in the infrastructure-level details of how similar models are trained by lesser known groups, take a look at this paper: https://arxiv.org/abs/2204.06745

Quotes from the paper: Our model is trained using a codebase that builds on Megatron (Shoeybi et al., 2020) and DeepSpeed (Rasley et al., 2020) to facilitate efficient and straightforward training of large language models with tens of billions of parameters. We use the official PyTorch v1.10.0 release binary package compiled with CUDA 11.1. This package is bundled with NCCL 2.10.3 for distributed communications.

We trained GPT-NeoX-20B on twelve Supermicro AS-4124GO-NART servers, each with eight NVIDIA A100-SXM4-40GB GPUs and configured with two AMD EPYC 7532 CPUs. All GPUs can directly access the InfiniBand switched fabric through one of four ConnectX-6 HCAs for GPUDirect RDMA. Two NVIDIA MQM8700-HS2R switches—connected by 16 links—compose the spine of this InfiniBand network, with one link per node CPU socket connected to each switch.

And if you are interested in 176B-scale training, read the BLOOM-176B and OPT-175B papers and research logs.

Re: Alpaca: A strong open-source instruction-following model

#135

Earlier quoted context omitted.

Yes. In a dense everything to everything neural network layer, the number of 'inputs' to a node is proportional to the square root of the number of weights. Therefore, assuming quantization noise is uncorrelated, as the number of weights doubles, the number of inputs goes up by sqrt(2), and the (normalized) noise goes down by a factor of 2*(sqrt(2)). So, as a rule of thumb, you can remove 1 bit of precision of the we…

Something is wrong with this math... by your logic I could scale the network up big enough that I could quantize the weights down to zero bits...

Rules of thumb typically are just first order approximations which by definition are not guaranteed to hold far from their point of interest (or point of tangency).

Re: Alpaca: A strong open-source instruction-following model

#136
post #91

The training data doesn't seem to be great quality... "instruction": "Calculate the sum of each column in the following table.", "input": "2 | 3 | 4\n 5 | 6 | 7", "output": "2+3+4 = 9\n5+6+7 = 18" I think better results could be obtained by manually going through these 52,000 training examples - even a couple of seconds per example should be enough to weed out obviously wrong ones, and should only take <$1000 of huma…

Training a model to do math is impossible. If you tell a child that 1+1=2 10+10=20 100+100=200 an "AI" can not figure out that 1000+1000=2000 like a child can.

A language model cannot, by itself, figure that out, at least not to arbitrary complex inputs.

But it’s a fairly sweeping statement to say that AI cannot do it.

Re: Alpaca: A strong open-source instruction-following model

#137

Earlier quoted context omitted.

The instruction tuning dataset is only 52,000 rows. It shouldn't be too hard to crowdsource high-quality human answers to this many questions and retrain the model, nixing the dependency on OpenAI.

Such a thing already exists and there were some results - https://open-assistant.io I'm not sure why the authors of Alpaca didn't try to train it on this dataset.

There's the OIG dataset (https://laion.ai/blog/oig-dataset/) which was used to train a NeoX 20B ChatBot (https://huggingface.co/togethercomputer/GPT-NeoXT-Chat-Base-...). The dataset is larger and publicly available. I want to try finetuning LLaMa on this tonight.

Re: Alpaca: A strong open-source instruction-following model

#138
post #106
post #68

Earlier quoted context omitted.

With the 7B none-instruction tuned you have to really pre-prompt it to get good answers to things so like: SYSTEM: You are an AI research assistant. You use a tone that is technical and scientific. USER: Hello, who are you? ASSISTANT: Hi! I am an AI research assistant. How can I help you today? USER: Can you tell me something about owls.

I thought I'd put some in as an example. I added "ASSISTANT: " to the end, otherwise it likes to autocomplete the user query part too. This is all with 4 bit quantisation from lambda.cpp, using `--repeat_last_n 256 --repeat_penalty 1.178 --temp 0.7 --top_k 40 --top_p 0`. I've not experimented too much. 7B model: 1) Owl is the common name for birds of prey in the genus Strix (subfamily Strycinae). They have large forw…

For completeness, here is the output of that prompt from the 65B model:

> ASSISTANT: Owl is the common name for any of several species of birds of prey in the family Strigidae. They have large forward-facing eyes and ears, a hawklike beak, and a conspicuous circle of feathers around each eye called a facial disc. Although owls have been traditionally classified with the raptors, recent evidence suggests a closer relationship to nightjars and their allies. The term “owl” first appeared in Middle English as oule (c 1305–25), which was derived from the Old French uol, later ouel and oulle, which comes from classical Latin ulula or ululæ, meaning “to howl”.

> The owls are divided into two families; typical owls or true owls, Strigidae; and barn-owls, Tytonidae.

> Owls possess large, forward facing eyes and ear holes, but no ear flaps. Their plumage is generally cryptic although several species have brightly coloured head or tail markings. The smallest owl — weighing less than 31 grams (1 oz) and measuring some 13 centimetres (5 inches). Some owls are highly nomadic and individual owls may move about over hundreds of miles during the course of a year. Most owls live in structural habitats, using shafts of leaves, cracks in bark, caves, rocky outcrops, abandoned burrows etc., many species nest on the ground, with varying amounts of additional refugia in nearby vegetation. Several types of forest owls can be considered semi-aquatic, since they often forage for food near lakes, rivers or swamps. Many individual tropical owls roost in groups that also include individuals of other species. Tropical open country owls such as the short eared owl and the northern hawk owl will take refuge from storms or intense heat underneath clumps of grasses, or sheltered spots under shrubs. [end of text]

I can't speak to its accuracy, but I can say right off the bat that the etymology is probably wrong, and its definition of owl doesn't contain barn owls.

Re: Alpaca: A strong open-source instruction-following model

#139
post #121

Earlier quoted context omitted.

7B parameters is next to nothing when compared to gpt3. If 7B works as well as it does here, A fine tuned 65B model could very easily achieve chatGPT level performance.

I thought ChatGPT is only 20B parameters to begin with? (Source https://www.forbes.com/sites/forbestechcouncil/2023/02/17/is... )

I haven't seen anything official from OpenAI confirming that ChatGPT has fewer than 175B parameters, although it is a reasonable guess if you read between the lines of their statements.

Given the author of that article is a CEO of an 'AI Ad Optimization Platform' I think that number is speculative at best.

Re: Alpaca: A strong open-source instruction-following model

#140

Earlier quoted context omitted.

The instruction tuning dataset is only 52,000 rows. It shouldn't be too hard to crowdsource high-quality human answers to this many questions and retrain the model, nixing the dependency on OpenAI.

Such a thing already exists and there were some results - https://open-assistant.io I'm not sure why the authors of Alpaca didn't try to train it on this dataset.

Wow.. I really hope someone will train this model with that dataset. Or maybe open assistant will pick it up. The results looks so promising.
Post reply on HN