Viewing profile — cgorlla
cgorlla
HN member- Joined
- Wed, Apr 29, 2020, 10:00 PM UTC
- HN karma
- 135
- Public activity
- 38 items
- HN profile
- View on Hacker News ↗
About cgorlla
No profile information was provided.
Recent public activity
-
comment
Comment #49199563
Based on the reception of our last post we took folks' suggestions to run the new official build of V4 Flash and compare it to the preview that was released 5 days apart. The new b…
- story
-
comment
Comment #49190389
Every post is actually scanned for LLM content as well.
-
comment
Comment #49190248
A typographic Voight-Kampff test is pretty awesome.
-
comment
Comment #49187545
[dead]
- story
-
comment
Comment #49117376
We discuss this in the writeup. While we expected this result, it is important for there to be data backing the claims, and an experimental setup that mirrors productions tasks is …
-
comment
Comment #49117328
We're actually exploring the changes in the model geometry that cause it to comply or not comply with a given policy next, I think visual representations of that behavior would be …
-
comment
Comment #49117298
I guess the AI that wrote your comment for you also conflated the SFT step of the target domain with the political prompts, which, in the sentence you quoted, contradicts your orig…
-
comment
Comment #49116748
>We plan to test what happens with a Chinese teacher into a Chinese-lineage base like Qwen next. :)
-
comment
Comment #49116739
The fact that you literally thought the examples were used in SFT in your last comment ago calls into question the utility of this conversation, notwithstanding the implication tha…
-
comment
Comment #49115938
The examples you're talking about are not involved in the training process, so their number is irrelevant. As stated in the post, the goal of this work is to determine whether a te…
-
comment
Comment #49115850
You can try it yourself! https://playground.ctgt.ai
-
comment
Comment #49115494
This is fixed
-
comment
Comment #49115488
This is fixed.
-
comment
Comment #49115476
Agreed, we find this to be an interesting reflection of societal values and norms inasmuch LLMs are.
- comment
- comment
-
comment
Comment #49115418
Abliterated models certainly have their uses but they're not the default choice for most users or enterprises, and thus not the versions of those models most would interact with.
-
comment
Comment #49115394
Agreed. It's fixed
-
comment
Comment #49115153
You can see exactly what prompts we used and the results here: https://github.com/CTGT-Inc/lineage-eval/tree/main/data We found V4 Flash was significantly more censored than the ba…
-
comment
Comment #49115010
It's most likely to occur when distilling a Chinese model from a Chinese base. We plan to do compliance geometry analysis in the future to see what is structurally changing in the …
-
comment
Comment #49114920
Consider that LLMs are trained on the corpus of the internet, and (simplifying) consequently give the average answer of the internet. If the desired answer of the censorer is contr…
- comment
-
story
Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
We recently used DeepSeek V4 Flash as a teacher for finance tasks with GPT-OSS-120B. Distillation works well on this problem. At a constrained 8k token budget, our self-distilled 1…