Live data from Hacker News

Viewing profile — sumo43

sumo43

HN member
Joined
Sat, Sep 09, 2023, 4:29 PM UTC
HN karma
216
Public activity
13 items

About sumo43

Interested in LLMs that browse the web natively. you can reach me at artem@lmresearch.net

Recent public activity

  1. comment
    Comment #45792294

    Try running this using their harness https://huggingface.co/flashresearch/FlashResearch-4B-Thinki...

  2. comment
    Comment #45792273

    I made a 4B Qwen3 distill of this model (and a synthetic dataset created with it) a while back. Both can be found here: https://huggingface.co/flashresearch

  3. story
  4. comment
    Comment #42012168

    I think the fine tuned policies are still very brittle, but I agree that this is super promising. It's also one of the most open (the model is still closed) research blogposts we'v…

  5. story
  6. comment
    Comment #40469265

    seems like an improvement on the aloha approach? You still need to finetune it on roughly the same amount of OOD examples. Contrast this with google's approach over 2023, which was…

  7. comment
    Comment #40229861

    Location: US Remote: Yes Willing to relocate: Yes (US) Technologies: Python, PyTorch, HuggingFace, C++ Résumé/CV: https://drive.google.com/file/d/1qY-m1tKz4_QpHgxaGryC2vk-DGs... Em…

  8. story
  9. comment
    Comment #39843063

    Maybe true for instruct, but pretraining datasets do not usually contain GPT-4 outputs. So the base model does not rely on GPT-4 in any way.

  10. comment
    Comment #39218440

    SEEKING VOLUNTEERS: open source self-play training for language models we are a small team associated with EleutherAI. looking to push the frontier of open source language models t…

  11. comment
    Comment #37560963

    For training you would need more memory. As for the pooling, Theoretically yes but wouldn't latency play as much, if not a greater part in the response time here? Imagine a tensor-…

  12. comment
    Comment #37551369

    Cool service. It's worth noting that, with quantization/QLORA, models as big as llama2-70b can be run on consumer hardware (2xRTX 3090) at acceptable speeds (~20t/s) using framewor…

  13. comment
    Comment #37447193

    Hello, I'm planning to participate in this challenge. I have experience training/prompting and building products from LLMs, I've also participated in a few CTFs. sumo43@proton.me