Live data from Hacker News

Will It Mythos?

swelljoe.com

161–170 of 232 posts

Re: Will It Mythos?

#161
post #77

As I posted in another comment, I found Fable to be substantially more powerful than any previous model. However, this isn't just an ungrounded opinion - I uploaded my full session transcript and code created working on a very complex implementation, so people can judge for themselves, if they're interested: https://tossrock.substack.com/p/36-hours-with-fable

> code created working on a very complex implementation I always find it amusing when people claim "a very complex implementation". Sometimes it's a hard problem, other times an easy one. Either way that's not for you to judge. And the implementation being complex... is that a good thing? Wouldn't a simple implementation be better? It reminded me of the parable of two programmers.

>> Either way that's not for you to judge.

Says who? If you find something complex, you can just say that it's complex. I don't get what the objection is.

Re: Will It Mythos?

#162

[flagged]

> there is no 'profit' step.

You have to learn to think like a drug dealer. The first hit is always free.

Companies and developers are growing more and more dependent on coding agents. Eventually, the owners of the AI will be able to charge whatever they want. What are you going to do? Go back to coding by hand? Do you even remember how?

Re: Will It Mythos?

#163

Earlier quoted context omitted.

Early on, I had a vague suspicion that the reason some of the Chinese models, including quite small ones, perform so well on this task, especially relative to their size and cost, is because they don't have the same safety guardrails baked in regarding software security that US models seem to have. Gemini 3.1 Pro doing so poorly sort of reinforced that gut feeling. But, then Gemma 4 proved to be extraordinarily good…

I concur with "Gemma 4 31B the best model I have results for". My workflow includes a lot of Gemma 4 – but dense 31B non-quantised version.(BTW I found it is most cost effective to run on Bedrock)

I tried to prove quantization made models worse, but in my testing Qwen 3.6 27b performed statistically the same from 4 bits to 16, using the unsloth dynamic quantizations. Gemma 4 4-bit QAT seems to perform the same as the full-fat version, but quite a lot faster.

But, I have come to consider Gemma 4 31b the best model I can self-host, even though there are bigger models that'll fit on the Strix Halo. (I could also use much bigger MoE models on my desktop which has 64GB VRAM and 112GB system RAM.)

Re: Will It Mythos?

#164

Earlier quoted context omitted.

At the time a GPT subscription didn't include Pro usage in the rolling limits. It was billed at API rates. Does it now? If anyone wants to fund the other five cases (~$125), I'll run them. I find that an unrealistic cost, though...simply not useful data. I'm certainly not going to spend $23 per file to audit a project with hundreds or thousands of files. I don't know anyone who would. Also note that it was $100 cap p…

I have ~100$/mo sub and I have Pro in chat app and Extra High in Codex for GPT-5.5 I think on sub tokens might be 100 times cheaper. The quota is also generous in my opinion. I can vibecode a lot most days of the week and not run out.

But GPT 5.5 on extra high is not Pro. When I looked into it, Pro was not available for agentic use via any rolling limits plan. But, I'll look again into whether there's some reasonable way to complete the test for GPT Pro.

Re: Will It Mythos?

#165
post #84

Around February, Opus 4.6 was excellent. Smart, fast, proactive. Then it got lobotomized and it's never been the same after that nerf. 4.7 came along and it too was disappointing—not unlike 4.8, which despite feeling a smidge smarter, tends to write word salad and is basically unusable for some workflows. Fable felt like having access to that "old Opus" again, but a little smarter. Sort of like I'd expect an Opus 5 t…

This is exactly what I find frustrating. I get comfortable with the latest model X. Then a new sparkly model Y launches. I am like, I don't need your new fangled Y, that consumes more tokens. My needs are small and i am happy with the older X. But then X starts to degrade. At first subtly, and then drastically. So then I am forced to upgrade to Y. What I do not understand is: > is this a sneaky way for companies to p…

[deleted]

Re: Will It Mythos?

#166
post #84

Around February, Opus 4.6 was excellent. Smart, fast, proactive. Then it got lobotomized and it's never been the same after that nerf. 4.7 came along and it too was disappointing—not unlike 4.8, which despite feeling a smidge smarter, tends to write word salad and is basically unusable for some workflows. Fable felt like having access to that "old Opus" again, but a little smarter. Sort of like I'd expect an Opus 5 t…

This is exactly what I find frustrating. I get comfortable with the latest model X. Then a new sparkly model Y launches. I am like, I don't need your new fangled Y, that consumes more tokens. My needs are small and i am happy with the older X. But then X starts to degrade. At first subtly, and then drastically. So then I am forced to upgrade to Y. What I do not understand is: > is this a sneaky way for companies to p…

The economics of AI fall apart if you stay with the old model forever. No need to buy new GPUs or build new data centers.

Re: Will It Mythos?

#167
post #84

Earlier quoted context omitted.

This is exactly what I find frustrating. I get comfortable with the latest model X. Then a new sparkly model Y launches. I am like, I don't need your new fangled Y, that consumes more tokens. My needs are small and i am happy with the older X. But then X starts to degrade. At first subtly, and then drastically. So then I am forced to upgrade to Y. What I do not understand is: > is this a sneaky way for companies to p…

The economics of AI fall apart if you stay with the old model forever. No need to buy new GPUs or build new data centers.

[dead]

Re: Will It Mythos?

#168

As I posted in another comment, I found Fable to be substantially more powerful than any previous model. However, this isn't just an ungrounded opinion - I uploaded my full session transcript and code created working on a very complex implementation, so people can judge for themselves, if they're interested: https://tossrock.substack.com/p/36-hours-with-fable

What tool did you use to export the transcript as HTML?

Re: Will It Mythos?

#169
post #65

As I posted in another comment, I found Fable to be substantially more powerful than any previous model. However, this isn't just an ungrounded opinion - I uploaded my full session transcript and code created working on a very complex implementation, so people can judge for themselves, if they're interested: https://tossrock.substack.com/p/36-hours-with-fable

Interesting. I tried Fable vs Codex 5.5 xhigh on three different cases. 1. A resource leak with unknown cause. Both of them zoomed onto the same potential issue and proposed almost identical patches. Fable missed an edge case that Codex handled correctly. 2. Review of a SPICE model. Models had different comments, none substantial. Both missed important issues that were simulated inadequately. Clearly a valley where t…

Did you use their native harnesses, or a generic one?

Re: Will It Mythos?

#170

Around February, Opus 4.6 was excellent. Smart, fast, proactive. Then it got lobotomized and it's never been the same after that nerf. 4.7 came along and it too was disappointing—not unlike 4.8, which despite feeling a smidge smarter, tends to write word salad and is basically unusable for some workflows. Fable felt like having access to that "old Opus" again, but a little smarter. Sort of like I'd expect an Opus 5 t…

All of these discussions of models being "nerfed" reminds me of discussions among audiophiles "this cable sounds so much better than this other one, it's night and day, ferrari versus honda civic" Yet when you do blind tests they can't tell the difference between a $1000 cable and a $1 one. I bet if you do blind tests between GPT-5.3, 5.4 and 5.5 most would struggle to tell them apart, yet they are certain that "5.5…

That's a pretty shallow dismissal, and I bet you $100 I can tell you which model I'm talking to between 4.6 and 4.8 without looking or asking after a handful of messages.

Anthropic famously had a terrible outage back when 4.6 was the latest and greatest, and it was never the same after it came back.

All evidence suggests they simply don't have the compute to keep serving their best models at their most powerful.

Post reply on HN