Live data from Hacker News

Viewing profile — behohippy

behohippy

HN member
Joined
Thu, Feb 18, 2016, 6:22 PM UTC
HN karma
200
Public activity
52 items

About behohippy

No profile information was provided.

Recent public activity

  1. comment
    Comment #48987279

    It's pretty simple nowadays if you know conceptually how they work. Running the LLM calls in a loop with tools is an agent. You only need 10 or so basic tools to accomplish nearly …

  2. comment
    Comment #48007363

    You might have a business idea there. I wouldn't mind a twinscan plushie for sitting on top of the workstation.

  3. comment
    Comment #44865507

    Yeah 48g, sub 200W seems like a sweet spot for a single card setup. Then you can stack as deep as you want to get the size of model you want for whatever you want to pay for the po…

  4. comment
    Comment #44865473

    Sure, all the slop code projects I produce get MIT licensed on public repos. It wasn't mine to begin with, so I wouldn't prevent anyone from using it.

  5. comment
    Comment #44832309

    Used 3090s have been getting expensive in some markets. Another option is dual 5060ti 16 gig. Mine are lower powered, single 8 pin power, so they max out around 180W. With that I'm…

  6. comment
    Comment #44119119

    About 768 gigs of ddr5 RAM in a dual socket server board with 12 channel memory and an extra 16 gig or better GPU for prompt processing. It's a few grand just to run this thing at …

  7. comment
    Comment #43492541

    These articles are gold, thank you. I used your gemma one from a few weeks back to get gemma 3 performing properly. I know you guys are all GPU but do you do any testing on CPU/GPU…

  8. comment
    Comment #43039504

    I run the KV cache at Q8 even on that model. Is it not working well for you?

  9. comment
    Comment #43024026

    Qwen is a little fussy about the sampler settings, but it does run well quantized. If you were getting infinite repetition loops, try dropping the top_p a bit. I think qwen likes l…

  10. comment
    Comment #43012090

    You probably won't be running fp16 anything locally. We typically run Q5 or Q6 quants to maximize the size of the model and context length we can run with the VRAM we have availabl…

  11. comment
    Comment #42792165

    Just this pic: https://imgur.com/ip8GWIh

  12. comment
    Comment #42792159

    I don't have a video but here's a pic of the output: https://imgur.com/ip8GWIh

  13. comment
    Comment #42792152

    It's a 3b model so the creativity is pretty limited. What helped for me was prompting for specific stories in specific styles. I have a python script that randomizes the prompt and…

  14. comment
    Comment #42785105

    I have a mini PC with an n100 CPU connected to a small 7" monitor sitting on my desk, under the regular PC. I have llama 3b (q4) generating endless stories in different genres and …

  15. comment
    Comment #40133129

    I had this same issue with incomplete answers on longer summarization tasks. If you ask it to "go on" it will produce a better completion, but I haven't seen this behaviour in any …

  16. comment
    Comment #38890822

    It's probably an evolution of the phi-1/1.5 "Textbooks are all you Need" training method: https://arxiv.org/abs/2309.05463

  17. comment
    Comment #38487772

    No joke, that would be an awesome LLM project name!

  18. comment
    Comment #38473955

    Top_p and top_k are pretty important concepts for LLMs same as temperature so P,K,C and F are underutilized

  19. comment
    Comment #36382660

    Hey emad, thanks for SD and this! What's the plan if Meta does Apache 2.0 for LLaMA? Just keep going and making the 30b and 65b or build different models?

  20. comment
    Comment #36217573

    Vicuna-13b (4bit) got the answer right, the first time as well.

  21. comment
    Comment #34803256

    I've noticed the same with my Asus TUF laptops. I've had 2 generations of the 15" models with Ryzen processors and adding a second stick seemed to "wake" them up in a noticeable wa…

  22. comment
    Comment #34352724

    My dad builds houses in northern Ontario and heat pumps seem to be getting more popular on new builds. This is a place that regularly gets below -30C in the winter. The heat pump (…

  23. comment
    Comment #34234891

    Not a good fit for lifestyle. I'm rural, and I haul around ATVs and dirt bikes with the main destination being 300km away and even more rural. My pickup truck is a way better choic…

  24. comment
    Comment #33733133

    I love this idea. I've been desperate enough to reuse gaskets in a few repairs, and suffer oil leaks down the line. Usually it was because of back ordered gaskets that would be wee…

  25. comment
    Comment #33343191

    Asus TUF series (A15). Inexpensive with really good hardware specs and upgradable ram, storage and wifi cards. I usually go AMD on them, and midrange video cards. They're shockingl…