Earlier quoted context omitted.
> Im afraid that the usual mantra that "we just need more scale" that worked well for attracting investments, is not working anymore - bigger models provide marginal improvements while naturally get much more expensive to run. It's super interesting to hear this refrain on HN, it is alarmingly common. Anthropic released benchmark numbers on Mythos, as they have for all of their models. Once models become public, peop…
Mythos numbers are effectively irreproducible aside from cherry-picked approvals.
Expanding Project Glasswing
231–240 of 261 posts
Re: Expanding Project Glasswing
#232Earlier quoted context omitted.
It is already very plausible (and has been since the 1950s) without the advent of LLMs. This is just another layer on top of the preexisting and very plausible existential threats we already face.
Detail it. Justify it. Your comment about before LLMs is a non sequitur. Demonstrate that an LLM can kill everyone on the planet.
There can be arms-races in domains that are unfathomable to the participants. A small mammal will die a billion times over before it understands the evolutionary mechanisms and the genetic playing field on which it loses. Actors are not necessarily privy to understand the means by which they will lose, and humans have only existed in a small window of time in which we fashioned a manicured garden, in which that full understanding was briefly possible. It is not favoured in the universe for us to fully understand our environment imho
If the risk must be exhaustively detailed before it is given credence, we are already doomed, and deservedly so
Re: Expanding Project Glasswing
#233Earlier quoted context omitted.
I had a geniunely surreal conversation with the security team the past week, it went like: 'Hi, we are reaching out to you because our tool flagged a large data transfer between such and such services' 'Wait, the source endpoint is an internal service, the target endpoint is an internal S3 bucket (I was doing a routine DB backup) Neither are reachable from the internet. How is it a security issue?' 'Our tool has flag…
Almost all the corporate security professionals I have dealt with have been tool runners with no more than Helpdesk level skills.
Re: Expanding Project Glasswing
#234Earlier quoted context omitted.
Detail it. Justify it. Your comment about before LLMs is a non sequitur. Demonstrate that an LLM can kill everyone on the planet.
Task a squirrel with justifying the risk of a fox, but from the biomolecular level. That is the level of the task you are setting out. There can be arms-races in domains that are unfathomable to the participants. A small mammal will die a billion times over before it understands the evolutionary mechanisms and the genetic playing field on which it loses. Actors are not necessarily privy to understand the means by whi…
Thats a really deep thought for a 12 year old.
>There can be arms-races in domains that are unfathomable to the participants.
You cant even justify LLMs as being unfathomable. Oh watch out I am fathoming them. You cant stop me fathoming all over the place.
>A small mammal will die a billion times over before it understands the evolutionary mechanisms and the genetic playing field on which it loses.Actors are not necessarily privy to understand the means by which they will lose, and humans have only existed in a small window of time in which we fashioned a manicured garden, in which that full understanding was briefly possible. It is not favoured in the universe for us to fully understand our environment imho
Non Sequitur. One that sounds like it was made up for that "What the Bleep" garbage.
>If the risk must be exhaustively detailed before it is given credence, we are already doomed, and deservedly so
The risk needs to be justified as something more substantial than weird people writing wannabe edgy messages on the internet. If someone on the internet told you that we need to drastically reverse living standards because there's a risk that modern technology will summon King Kong any reasonable person would ask for the working out instead of running for a cave.
Re: Expanding Project Glasswing
#235Earlier quoted context omitted.
Almost all the corporate security professionals I have dealt with have been tool runners with no more than Helpdesk level skills.
That means you aren't high enough up to deal with the non helpdesk level security people.
Re: Expanding Project Glasswing
#236Earlier quoted context omitted.
Task a squirrel with justifying the risk of a fox, but from the biomolecular level. That is the level of the task you are setting out. There can be arms-races in domains that are unfathomable to the participants. A small mammal will die a billion times over before it understands the evolutionary mechanisms and the genetic playing field on which it loses. Actors are not necessarily privy to understand the means by whi…
>Task a squirrel with justifying the risk of a fox, but from the biomolecular level. That is the level of the task you are setting out. Thats a really deep thought for a 12 year old. >There can be arms-races in domains that are unfathomable to the participants. You cant even justify LLMs as being unfathomable. Oh watch out I am fathoming them. You cant stop me fathoming all over the place. >A small mammal will die a…
Re: Expanding Project Glasswing
#237Earlier quoted context omitted.
The question is, will anyone pay enough for Mythos to offset the opportunity cost of offering that much Opus? You don't want to end up in a spot where you don't have enough compute and your service's reliability degrades to an unusable state like xAI.
I feel like there's always a demand for the very best models, even at insane prices. If the opportunity cost is x times opus, maybe few but there will always be companies willing to pay x+1 times opus.
Re: Expanding Project Glasswing
#238Earlier quoted context omitted.
>Task a squirrel with justifying the risk of a fox, but from the biomolecular level. That is the level of the task you are setting out. Thats a really deep thought for a 12 year old. >There can be arms-races in domains that are unfathomable to the participants. You cant even justify LLMs as being unfathomable. Oh watch out I am fathoming them. You cant stop me fathoming all over the place. >A small mammal will die a…
You're kind of an asshole. No thanks
Re: Expanding Project Glasswing
#239Earlier quoted context omitted.
As someone with over 30 years experience in computer security, both in corporate as well as boutique security and startup shops, who has been consistently fighting this trend, and recently bearing witness to and engaging in the current AI surge: I can say with absolute confidence that it is only getting and going to get even worse yet. People like me who know there is a better way are getting pushed harder to lean on…
AI is far better at security than the majority of security professionals. It is a net positive. People constantly compare AI to this very rare expert human rather than the reality of who is already employed. Experts like you are a major culprit of this. And it puts you at odds with yourself to both admit the industry is full of subpar workers and then lament that they will be replaced with workers that are better, bu…
We need experts to know when AI is wrong, which it is all the time.
Earlier this week someone commented here that we shouldn’t expect a language model to know that you need to drive a car to a car wash, to wash a car.
So then, what do we expect it to know? Who’s responsible for when it’s wrong?
Also, why can’t Mythos just fix all these issues itself if it’s so smart. And test them to make sure they work?
Re: Expanding Project Glasswing
#240Earlier quoted context omitted.
“Cybersecurity weather person and award winning shitposter.” why are they someone we should pay attention to the opinion of?
The most intelligent person involved in the highest level projects at my current company introduces themselves as an out of work circus clown. There is an incredible amount of competency signaled by someone who was given access to this model but doesn’t treat their online presence like a professional resume.