Live data from Hacker News

Making my local LLM voice assistant faster and more scalable with RAG

johnthenerd.com

1–10 of 17 posts

Re: Making my local LLM voice assistant faster and more scalable with RAG

#3
post #2

That lag between query and response ruins it for me.

“Excellent query good sir! …”

And more seriously, it seems like the LLM could be used to precreate lots of filler prefixes that correspond to the rag’d document that are being sent to the model.

While it wouldn’t work if you’re GPU’d bound, multiple prompts could be run in parallel with different pieces of context and then have the model chose the most appropriate response (which could be done in parallel too).

Re: Making my local LLM voice assistant faster and more scalable with RAG

#7

Cringe conversation. Why can't AIs just do stuff that you ask them to do without pretending to be human?

Llama3 is very keen to be nice. I kind of wonder if that's due to better results on the chatbot arena (probably not, just a conspiracy theory I like). But with enough context available, you can definitely tweak the response in many ways. Give an example or two, tell it to be an emotionally detached HAL, you'll get what you want.

Re: Making my local LLM voice assistant faster and more scalable with RAG

#9
I was having a look at the model mentioned, specifcially `casperhansen/llama-3-70b-instruct-awq`.

When checking this model, I found out [1] it's based on llama-2 ?

``` Expand Llama 3 70B Instruct AWQ Parameters and Internals LLM Name Llama 3 70B Instruct AWQ Repository Open on Base Model(s) Llama 2 70B Instruct quantumaikr/llama-2-70B-instruct Model Size 70b ```

I added a question [2] on Hugging Face to learn more about this.

Anyone could explain to me what this means? Does it mean that it has been trained on the version 2 and wrongly named version 3? Or is it something that is not well intended?

[1] https://llm.extractum.io/model/casperhansen%2Fllama-3-70b-in...

[2] https://huggingface.co/casperhansen/llama-3-70b-instruct-awq...

Re: Making my local LLM voice assistant faster and more scalable with RAG

#10
post #6

Cringe conversation. Why can't AIs just do stuff that you ask them to do without pretending to be human?

Then how is it different from Excel/Word and shell/python scripts?

it's much slower, gets things wrong, and insists on things that ain't so.
Post reply on HN