Live data from Hacker News

Viewing profile — thntk

thntk

HN member
Joined
Wed, Nov 01, 2017, 4:21 PM UTC
HN karma
32
Public activity
15 items

About thntk

thnbiz [at] gmail [dot] com

Recent public activity

  1. comment
    Comment #44808990

    The model architecture only uses and cites pre-2023 techniques from the GPT-2 and GPT-3 era. Probably they intentionally tried to use the most bare transformers architecture possib…

  2. comment
    Comment #42622108

    Anyone know if it can run training/fine-tuning and not just 4-bit inference? Does it support mixed precision training with either BF16 or FP16?

  3. comment
    Comment #41433642

    It sounds like your startup is doing great, good luck.

  4. comment
    Comment #41427066

    Have you tried hiring or rotating internal people just for solving specific (management) tasks? It is like the "just-in-time" style in Japanese corps.

  5. comment
    Comment #41425330

    It's not founder mode or manager mode. I think it is just about effective management. When starting up, the founder needs to (1) know what should be done, (2) be able to do it them…

  6. comment
    Comment #41310955

    Isn't it the technological cycle in IT? We have seen PC softwares in the 90s, then web apps, then mobile apps, and now AI services. Except that currently AI is immature, so compani…

  7. comment
    Comment #41059670

    Anyone know what caused the very big performance jump from Large1 to Large2 in just a few months? Besides, parameter redundancy seems evidenced. Front-tier models used to be 1.8T, …

  8. comment
    Comment #41054007

    When Zuck said spy can easily steal models, I wonder how much of it comes from experiences. I remember they struggled to train OPT not long ago. On a more serious note, I don't rea…

  9. comment
    Comment #41048540

    Correct me if I'm wrong, my impression is that 3.1 is a better fine-tuned variant of base 3.0 with extensive use of synthetic data.

  10. comment
    Comment #40993196

    I've seen such articles more and more recently. In the past, when people had a vague idea, they had to do research before writing. During this process, they often realized some fla…

  11. comment
    Comment #40992784

    We knew high quality data can help as evidenced by the \Phi models. However, this alone can never eliminate hallucination because data can never be both consistent and complete. Mo…

  12. comment
  13. comment
    Comment #40897490

    Besides the burden of knowledge and the flaws of academic funding, another factor to explain the slowdown in scientific progress is: the world has more things to be entertained wit…

  14. comment
    Comment #16562200

    Oh, that is so ridiculous that I have to come back and LOL.

  15. comment
    Comment #16519927

    What does porn have to do with free speech and privacy?