Live data from Hacker News

Viewing profile — mattnewton

mattnewton

HN member
Joined
Sun, Jun 16, 2013, 11:34 PM UTC
HN karma
13,248
Public activity
3,240 items

About mattnewton

matthewnewton.com

Recent public activity

  1. comment
    Comment #49260286

    I think efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute; and we are not going to run out of e…

  2. comment
    Comment #49157956

    If China keeps releasing LLMs with such permissive licenses it probably works better since the pretraining rnd is subsidized and de-risked - but that’s a big if.

  3. comment
    Comment #49157910

    The exception being the token embeddings and lm head (which scale with the number of tokens the model knows and presumably you need a smaller number in the tokenizer for only Engli…

  4. comment
    Comment #49151553

    Honestly the 27b dense one punches way above its weight in a lot of domains, especially coding in my testing, so I think you will probably be disappointed.

  5. comment
    Comment #49151547

    I agree. 27b dense really did seem like the sweet spot.

  6. comment
    Comment #49151023

    There was a 3.5 122B 10A release - https://huggingface.co/Qwen/Qwen3.5-122B-A10B

  7. comment
    Comment #49023588

    Because demand for inference tokens is above supply

  8. comment
    Comment #48851878

    This was a big concern for earlier models, but with modern CoT trained models they should be able to come to the conclusion entirely in the thinking trace.

  9. comment
    Comment #48812107

    I think a) the labs are releasing very fast and b) why would they implement the long tail of app features when they can effectively sell tokens to every user to write their own ver…

  10. comment
    Comment #48779394

    It’s not that poor people can’t afford a stamp it’s that they aren’t going to spend money on an automated service or stamp if there are other places to apply to that don’t require …

  11. comment
    Comment #48768606

    Why do your think Meta doesn’t make money? Their ad platforms are incredibly lucrative.

  12. comment
    Comment #48768537

    There are plenty of services to send mail form the internet for a small fee, so this will only discourage the most poor candidates and add friction for the best ones.

  13. comment
    Comment #48755991

    I’m saying there is basically no way to both make vlms able to understand the long tail of PDFs where the layout conveys information (like charts and tables) and to make it as toke…

  14. comment
    Comment #48755269

    Because PDFs are a nightmare of a format and the only thing that’s is reasonably guaranteed about them is they will render to an image that people can read, the parsing of which wi…

  15. comment
    Comment #48755244

    But then I close my laptop and it’s not running on the headless host anymore right

  16. comment
    Comment #48754171

    I read point 19 as Palantir’s goal being to import Chinese and Russian style surveillance, and the comment saying effectively “it can’t happen here” and “taking them less seriously…

  17. comment
    Comment #48748803

    How hot does the water need to be before you raise the alarm? I think there is value in pointing out trend lines and voicing opposition even if there are other countries that have …

  18. comment
    Comment #48689392

    In the US, I think we are being intentionally DDOS-ed. This strategy was laid out by Steve Bannon in the old frontline PBS interview where he called the media “the opposition party…

  19. comment
    Comment #48689349

    Yeah people are overdoing it, and not everyone is being motivated to do something by feeling mad, but maybe instead of throwing up your hands that you can’t care about soybean tari…

  20. comment
    Comment #48674300

    _in a polling place_ no less

  21. comment
    Comment #48661925

    Definitely encourage you to test the models. We tried to optimize for realistic focus and not over-sharpening, which leads to a "hyper" AI-look. It's hard to benchmark because peop…

  22. comment
    Comment #48661328

    Krea 2 Large (on the website and api) was trained with the FLUX 2 VAE, if you want to test it out and push realism. After working with both I think the flux VAE has a slight edge i…

  23. comment
    Comment #48661242

    You can find some links and details in the GitHub readme for finetuning / LoRA support. Ostiris, musubi tuner, fal and hugging face diffusers are all day-0 supported :) https://git…

  24. comment
    Comment #48646660

    Hi HN, we're releasing weights for our latest text to image model and publishing this writeup on how we trained it in quite a bit of depth. I hope there is something in the report …

  25. story