Live data from Hacker News

Viewing profile — orost

orost

HN member
Joined
Mon, Aug 06, 2018, 4:42 PM UTC
HN karma
194
Public activity
34 items

About orost

No profile information was provided.

Recent public activity

  1. comment
    Comment #40066544

    Let me clarify. Mixtral-8x22B-v0.1 was released a couple days ago. The "mixtral:8x22b" tag on ollama currently refers to it, so it's what you got when you did "ollama run mixtral:8…

  2. comment
    Comment #40066456

    Considering "mixtral:8x22b" on ollama was last updated yesterday, and Mixtral-8x22B-Instruct-v0.1 (the topic of this post) was released about 2 hours ago, they are not the same mod…

  3. comment
    Comment #40066340

    That's not the model this post is about. You used the base model, not trained for tasks. (The instruct model is probably not on ollama yet.)

  4. comment
    Comment #39889021

    Mistral 7B Instruct v0.2 and Mistral 7B v0.2 are different models. Judging by the title, I suspect OP meant to post about the latter, which was released a few days ago, but acciden…

  5. comment
    Comment #39282410

    A rocket on a typical orbital launch profile spends less than 60 seconds in air dense enough for jet engines to have good performance, so there is little to gain. Pegasus is an orb…

  6. comment
    Comment #39248099

    An air-breathing jet engine doesn't need to carry oxidizer, which in a rocket is most of the propellant weight. It also has access to unlimited reaction mass, so it can be much mor…

  7. comment
    Comment #38736509

    You can partially offload with some backends (e.g. llama.cpp and derivatives) but speed gains from that don't come in until it's mostly offloaded. I have 8GB VRAM and it's not enou…

  8. comment
    Comment #38139748

    A reactor that has never been turned on isn't a significant radiation hazard. It's the fission products that are hazardous, not the fuel, if it's never gone critical there are no f…

  9. comment
    Comment #37892939

    The bazarek is fun but in reality even less relevant that this post makes it out to be. Since people with real information cannot prove it and it takes zero effort to post fakes th…

  10. comment
    Comment #37409857

    Preparations for pad repairs and upgrades were well underway before the first flight - the question was not whether they'd be necessary, but how much and how soon. In particular if…

  11. comment
  12. comment
    Comment #37070137

    Anything with 64GB of memory will run a quantized 70B model. What else you need depends on what is acceptable speed for you. With a decent CPU but without any GPU assistance, expec…

  13. comment
    Comment #36833114

    The simulation is just so fake, almost everything that goes on is just decorative. There is a budget, but after the first 30 minutes you'll always be running an enormous surplus wi…

  14. comment
    Comment #36555092

    Yes, many, huggingface is full of chat-tuned LLaMA derivatives that are supposed to replicate its performance, and tools like text-generation-webui or kobold.cpp can be used to run…

  15. comment
    Comment #36377336

    Experimental Falcon inference via ggml (so on CPU): https://github.com/cmp-nct/ggllm.cpp It has problems but it does work

  16. comment
    Comment #36222597

    ggml is a library that provides operations for running machine learning models llama.cpp is a project that uses ggml to run LLaMA, a large language model (like GPT) by Meta whisper…

  17. comment
    Comment #36204844

    It doesn't matter very much that the "official" instruct tune is censored as anyone can create their own and there will probably be many freely available ones as happened with LLaM…

  18. comment
    Comment #36204519

    You have to turn down temperature and/or p when you want accuracy. Otherwise you don't know if the model's read is bad, or if you just happened to get a low-probability outlier. Wi…

  19. comment
    Comment #36139126

    You can just barely fit a 33B GPTQ model in 24GB VRAM. It will be in 4-bit mode, and without maximum context size, but it will be quite fast. Or you can run from RAM+VRAM in GGML f…

  20. comment
    Comment #36065805

    Quantization isn't (and wasn't) expensive, it's mostly just data shuffling. A good PC will do a 7B model in half a minute, up to a few minutes for a larger model. Quantized models …

  21. comment
    Comment #35908753

    Almost every UI for LLMs I've seen has a way to specify an initial prompt that never goes out of context, it's strange that it's not a feature in ChatGPT.

  22. comment
    Comment #35851583

    I definitely noticed a drop in quality when the gimped (but presumably dramatically cheaper to run) GPT-3.5-turbo model was introduced on the free version. As a paying subscriber I…

  23. comment
    Comment #35772753

    The charges worked fine. They blew holes in the tanks as planned. The issue is that this didn't cause immediate structural failure as intended. But structural failure of the rocket…

  24. comment
    Comment #35766243

    There is nothing out there that quite matches ChatGPT quality but you can get a similar kind of experience by running an instruction-tuned derivative of LLaMA with llama.cpp. Try s…

  25. comment
    Comment #35639559

    I suspect you could train a model to just shut up and follow instructions. I.e. instead of "Do X -> Sure! As a large language model, I'd love to help you with X!", just "Do X -> X"…