Live data from Hacker News

Will It Mythos?

swelljoe.com

101–110 of 232 posts

Re: Will It Mythos?

#101
I find this interesting:

  …no model performed better with an Agent, a couple performed worse, and time/tokens/costs were consistently much higher with the agent in the loop, for some reason.
Somone should build a harness where features are only added if they are proven net positive to outcomes.

Re: Will It Mythos?

#102
post #82

Earlier quoted context omitted.

For reference: if you write the code as cleverly as possible, you are, by definition, not smart enough to debug it

Ah thanks - I couldn't remember the original version. For reference: it's called Kernighan's Law, and can be found in the Second Edition of "The Elements of Programming Style", page 10 [1]. The original phrasing is: > Everyone knows that debugging is twice as hard as writing a program in the first place. So if you’re as clever as you can be when you write it, how will you ever debug it? [1] https://archive.org/detail…

It seems I was not able to either, and I trusted google AI snippet. Thanks

Re: Will It Mythos?

#103
post #98

I've read opinions that this a speculation to raise the Anthropic's value. They are known to say "horrific things" and personification of the AI they are delivering. It sometimes sounds unprofessional even. This line of communication might have even influenced the courts in the case of copyright violation ("it is not copyright violation if a person learned something and it knows it and thinks of it"). However algorit…

> They are known to say "horrific things" and personification of the AI they are delivering. It sometimes sounds unprofessional even.

It doesn’t sound unprofessional— it sounds unethical. Either they’re making something that they genuinely believe is unsafe but don’t want to stop because, you know, that’s business! Have you seen how much this shit costs? Or they’re deliberately making the entire country feel unsafe because it looks great to investors. Either way, frankly, fuck them and everybody else playing this dumb billionaire’s game. They deserve every bit of static this dimwitted government levels at them.

Re: Will It Mythos?

#104

Fable was the only model that was able to detect a data corruption bug in my Qt C++ note-taking app[1] that all other tested models (gpt-5.5 xhigh, GLM-5.1, Kimi 2.7, DeepSeek V4 Pro) didn't find. I'll test on GLM-5.2 and Mimo v2.5 Pro soon. [1] https://www.get-notes.com

I asked Fable on max to create a mathematical model to show that c (speed of light) is emergent from pregeometric physics.

It said: I can't, but it would be lazy to say that is is not a possibility.

With some back and forth it created a 5 step plan to narrow down if our universe has all the right properties for this to be true.

We evaluated the first four stages to be true, and it wrote the solver to find out if the fifth test running the full model passes, but that will take thousands of hours of compute.

Re: Will It Mythos?

#105

Around February, Opus 4.6 was excellent. Smart, fast, proactive. Then it got lobotomized and it's never been the same after that nerf. 4.7 came along and it too was disappointing—not unlike 4.8, which despite feeling a smidge smarter, tends to write word salad and is basically unusable for some workflows. Fable felt like having access to that "old Opus" again, but a little smarter. Sort of like I'd expect an Opus 5 t…

I miss the old Opus 4.6 too. They're probably quantizing the old models.

K/V cache compression and context shortening / summarisation. And yes, I suspected Quants too.

Re: Will It Mythos?

#106

Around February, Opus 4.6 was excellent. Smart, fast, proactive. Then it got lobotomized and it's never been the same after that nerf. 4.7 came along and it too was disappointing—not unlike 4.8, which despite feeling a smidge smarter, tends to write word salad and is basically unusable for some workflows. Fable felt like having access to that "old Opus" again, but a little smarter. Sort of like I'd expect an Opus 5 t…

All of these discussions of models being "nerfed" reminds me of discussions among audiophiles "this cable sounds so much better than this other one, it's night and day, ferrari versus honda civic" Yet when you do blind tests they can't tell the difference between a $1000 cable and a $1 one. I bet if you do blind tests between GPT-5.3, 5.4 and 5.5 most would struggle to tell them apart, yet they are certain that "5.5…

You will be amused to hear that when Anthropic "refreshed" 4.6 on AWS Bedrock I found it in my tests and wrote about it – and they actually rolled it back. This is how much non–coding tests may tell you about the model.

Re: Will It Mythos?

#107
> Note GPT 5.5 Pro is at the top of the leaderboard only because it blew through $100 budget after only completing four cases, so 2/4 is 50%. And, a couple of other results, both Qwen models, are skewed upward in the detect % ranking because of failure to complete all cases.

Try a Wilson score interval on the lower bound of the binomial proportion confidence interval [1].

So GPT 5.5 Pro’s 2/4 (p = 0.5) for one-sided 95% (z ~ 1.645), adjusts to 0.182 [a], and the top models are revealed as the 4/9s (mimo-v2.5-pro, gpt-5.5, opus-4.8, gemini-3.5-flash and deepseek-v4). (We need to dial CI down to 76% for gpt-4.5-pro to regain top status.) If we account for speed in that cohort, derpseek-v4 (91s) is fastest followed by opus-4.8 (137s).

Given deepseek-v4 is also the cheapest model among those five, I would say—based on these data—it’s the winner. (Out of the table. If Fable got 9/9, it’s obviously first.)

[1] https://en.wikipedia.org/wiki/Binomial_proportion_confidence...

Re: Will It Mythos?

#108
post #98

I've read opinions that this a speculation to raise the Anthropic's value. They are known to say "horrific things" and personification of the AI they are delivering. It sometimes sounds unprofessional even. This line of communication might have even influenced the courts in the case of copyright violation ("it is not copyright violation if a person learned something and it knows it and thinks of it"). However algorit…

> They are known to say "horrific things" and personification of the AI they are delivering. It sometimes sounds unprofessional even. It doesn’t sound unprofessional— it sounds unethical. Either they’re making something that they genuinely believe is unsafe but don’t want to stop because, you know, that’s business! Have you seen how much this shit costs? Or they’re deliberately making the entire country feel unsafe b…

Unless you think someone's going to build it, and either it's you or them, and you hope you can do it less horrifically.

Re: Will It Mythos?

#109
post #61

Earlier quoted context omitted.

> agency and misalignment are two sides of the same coin. The free will coin?

In my experience "free will", like "consciousness" and "common sense", is not so much a concept with a universally agreed definition as it is a cognitive stop sign or an applause light, meaning different things to everyone who uses the term. Do I have free will, or am I bounded by the laws of physics? Even if you think my soul is completely independent of my body, there are theologians who argue that God being omnisc…

Of all of the concepts like "consciousness" and "agency", "free will" is probably the least useful and poorly defined.

It's a hand-me-down from Western beliefs about morality and individuality - including Thelema and Christianity.

So there's a lot of starting from the concept and working back to assumed conclusions.

Generally humans do not have free will, do have very limited political, economic, and psychological agency, usually selected from a small number of competing rule sets, and are also far more easily influenced than they suspect.

Culture is more like a cellular automaton or diffusion system. Occasionally a transformation ripples out from an individual cell, often for fairly random reasons, but the big patterns are emergent, and every so often the soup shakes itself up and settles into a new arrangement.

IMO LLMs are the most recent proto-version of that, running on a different substrate.

Re: Will It Mythos?

#110

Earlier quoted context omitted.

Bugs the other models were benchmarked on are from the corpus that Mythos found. So Mythos might have 100% in this benchmark. Although the benchmark had 100$ budget cap and rudimentary tooling so probably a bit less than 100%. GPT-5.5-pro attemted only 4 problems out of 9 before the budget ran out and got 2 of them right. It's a shame that the author didn't try GPT-5.5-pro on all 9 just for completeness, pehaps on su…

At the time a GPT subscription didn't include Pro usage in the rolling limits. It was billed at API rates. Does it now? If anyone wants to fund the other five cases (~$125), I'll run them. I find that an unrealistic cost, though...simply not useful data. I'm certainly not going to spend $23 per file to audit a project with hundreds or thousands of files. I don't know anyone who would. Also note that it was $100 cap p…

I have ~100$/mo sub and I have Pro in chat app and Extra High in Codex for GPT-5.5

I think on sub tokens might be 100 times cheaper.

The quota is also generous in my opinion. I can vibecode a lot most days of the week and not run out.

Post reply on HN