Viewing profile — milkshakes
milkshakes
HN member- Joined
- Thu, Mar 05, 2009, 12:28 AM UTC
- HN karma
- 3,573
- Public activity
- 732 items
- HN profile
- View on Hacker News ↗
About milkshakes
Recent public activity
-
comment
Comment #49217734
you can't downvote a submission it's just that nobody upvoted it
-
comment
Comment #49216666
> models are still relatively easy to monitor and constrain are they?
- story
- story
- story
-
comment
Comment #49105820
they also have a pretty handy button right next to the error to ftfy
-
comment
Comment #49066558
i guess what i'm curious about is where you would draw the line, and why. how do you define a user agent? is it desirable or not for users to be able to discover and access resourc…
-
comment
Comment #49063958
i guess the thing i'm most confused about is what is the higher level goal here. 1 in 8 humans on the planet are experiencing the "web" through chatgpt alone. many have migrated to…
-
comment
Comment #49053871
start, not finish
-
comment
Comment #49053533
http://www.incompleteideas.net/IncIdeas/BitterLesson.html > One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that c…
-
comment
Comment #49047293
start here: https://huggingface.co/blog/mlabonne/abliteration
-
comment
Comment #49014644
what top comment?
-
comment
Comment #49008378
nyt is no fan of AI, but they cover this: https://www.nytimes.com/2026/07/09/business/china-russia-ai-...
-
comment
Comment #48999765
this is quite literally reward hacking. the model, under evaluation with cyber capabilities enabled, used those capabilities to simply bypass the exercise entirely and aim straight…
-
comment
Comment #48965153
who are you talking about? again, my question is concerned with the "second class labs" and the sustainability of the distillation-as-a-service model.
-
comment
Comment #48963130
it is not a required first step for training a model, sure. but that's not what i claimed. what i claimed is that is how they are so significantly _reducing the cost_ of training o…
-
comment
Comment #48963065
the obvious difference is the massive scale of data and compute required to develop and evolve these models, and the costs they impose on those building them.
-
comment
Comment #48962496
https://www.anthropic.com/news/detecting-and-preventing-dist... Moonshot AI Scale: Over 3.4 million exchanges The operation targeted: Agentic reasoning and tool use Coding and data…
-
comment
Comment #48962141
i never assumed that, and i do keep up with the publications. i'm also not saying it's a dumb thing to do! what i am saying is that empirically, it appears that distillation of a m…
-
comment
Comment #48962070
the question was: what is the endgame for the stated "second class labs" strategy of distilling their frontier competitors then undercutting them on price?
-
comment
Comment #48961941
the USG/NSA will fund chinese labs? to what end?
-
comment
Comment #48961740
> with budgets and what will fund these budgets exactly? inference is cheap, distillation is cheap, training is what's expensive.
-
comment
Comment #48961676
assume you are a "second class lab" and you are in fact making progress by distilling the results of the frontier labs' efforts. what is the end game for this strategy? if the fron…
-
comment
Comment #48952304
is this a joke?
-
comment
Comment #48899594
in fact, telegram does support e2e encryption ("secret chats")