Live data from Hacker News

Viewing profile — cgorlla

cgorlla

HN member
Joined
Wed, Apr 29, 2020, 10:00 PM UTC
HN karma
135
Public activity
38 items

About cgorlla

No profile information was provided.

Recent public activity

  1. comment
    Comment #49199563

    Based on the reception of our last post we took folks' suggestions to run the new official build of V4 Flash and compare it to the preview that was released 5 days apart. The new b…

  2. story
  3. comment
    Comment #49190389

    Every post is actually scanned for LLM content as well.

  4. comment
    Comment #49190248

    A typographic Voight-Kampff test is pretty awesome.

  5. comment
  6. story
  7. comment
    Comment #49117376

    We discuss this in the writeup. While we expected this result, it is important for there to be data backing the claims, and an experimental setup that mirrors productions tasks is …

  8. comment
    Comment #49117328

    We're actually exploring the changes in the model geometry that cause it to comply or not comply with a given policy next, I think visual representations of that behavior would be …

  9. comment
    Comment #49117298

    I guess the AI that wrote your comment for you also conflated the SFT step of the target domain with the political prompts, which, in the sentence you quoted, contradicts your orig…

  10. comment
    Comment #49116748

    >We plan to test what happens with a Chinese teacher into a Chinese-lineage base like Qwen next. :)

  11. comment
    Comment #49116739

    The fact that you literally thought the examples were used in SFT in your last comment ago calls into question the utility of this conversation, notwithstanding the implication tha…

  12. comment
    Comment #49115938

    The examples you're talking about are not involved in the training process, so their number is irrelevant. As stated in the post, the goal of this work is to determine whether a te…

  13. comment
    Comment #49115850

    You can try it yourself! https://playground.ctgt.ai

  14. comment
    Comment #49115494

    This is fixed

  15. comment
    Comment #49115488

    This is fixed.

  16. comment
    Comment #49115476

    Agreed, we find this to be an interesting reflection of societal values and norms inasmuch LLMs are.

  17. comment
  18. comment
  19. comment
    Comment #49115418

    Abliterated models certainly have their uses but they're not the default choice for most users or enterprises, and thus not the versions of those models most would interact with.

  20. comment
    Comment #49115394

    Agreed. It's fixed

  21. comment
    Comment #49115153

    You can see exactly what prompts we used and the results here: https://github.com/CTGT-Inc/lineage-eval/tree/main/data We found V4 Flash was significantly more censored than the ba…

  22. comment
    Comment #49115010

    It's most likely to occur when distilling a Chinese model from a Chinese base. We plan to do compliance geometry analysis in the future to see what is structurally changing in the …

  23. comment
    Comment #49114920

    Consider that LLMs are trained on the corpus of the internet, and (simplifying) consequently give the average answer of the internet. If the desired answer of the censorer is contr…

  24. comment
  25. story
    Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

    We recently used DeepSeek V4 Flash as a teacher for finance tasks with GPT-OSS-120B. Distillation works well on this problem. At a constrained 8k token budget, our self-distilled 1…