Live data from Hacker News

Fable and the end of the free lunch

dbreunig.com

231–240 of 268 posts

Re: Fable and the end of the free lunch

#231

Earlier quoted context omitted.

To be honest: I don't know for certain, but I'd assume that the people who pay the bills get the strongest alignment. They may not be tech people, but I (so far) haven't got a reason to think that the AI engineers are going behind the backs of their corporate leadership and subverting what they're being asked to do; do you? (I think it would be a good thing for humanity if they did)

I think right now both the engineers developing AI and the share holders are more focused on beating coding benchmarks and gaining revenue than anything to do with alignment.

That's alignment with shareholder value, at least.

Re: Fable and the end of the free lunch

#232
post #201
post #194

Earlier quoted context omitted.

I don't understand what this means. I use LLMs daily for my work in programming things, and they regularly will assert things that are not accurate.

When you let the the agent do a test, or tell it to read that doc first, you will ground it in reality. Doesn't mean they are 100% reliable. But without and on their own without access to grounding information, they halluzinate wildly.

The parent comment said that these types of behaviors are inherent to how LLMs work, and your response was to "ground them in facts". From what I can tell, this does not meaningfully change the original point the parent comment was making that the flaw is inherent; throwing a bunch of extra context at it to try to make it happen less often is useful, but it's still just a best-effort mitigation for the behavior, not somehow a way of literally changing the inherent nature of it.

Re: Fable and the end of the free lunch

#233
post #194

Earlier quoted context omitted.

I don't understand what this means. I use LLMs daily for my work in programming things, and they regularly will assert things that are not accurate.

A lot like humans, really. People regularly cite something they read, or quote a stat that turns out to be just completely inaccurate. But if you look up the thing, then you have facts again.

Sure, and by the same logic, if I was trying to cook a steak, I would not trust an arbitrary human to know the correct temperature off the top of their head; I'd want someone who I could trust had actual experience with the task I was trying to perform. The difference is that most humans are fully able to recognize whether they've cooked steak often enough to know the correct temperature off the top of their head, and they will say "I don't know" to most random arbitrary questions you ask them outside of their experience. I've yet to see an LLM product aimed at general usage for individuals be willing to say this without someone having to literally direct them to give that as an answer if they're not sure.

Re: Fable and the end of the free lunch

#234
post #211

Earlier quoted context omitted.

Giving Elon cash seems still wrong to me.

Worse than giving to Sama? (I mean that honestly.)

I think most would say yes? The sheer amount of misery Elon has caused with the DOGE antics would alone qualify, IMO.

Re: Fable and the end of the free lunch

#235
post #41
post #2

The real revolution is Deepseek v4 flash and similar models (GPT 5.6 Luna, muse spark 1.2, mimo, etc...) - Genuinely good performance for a tiny fraction of the cost of Fable and even GLM etc... I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis

I was using ChatGPT voice during cooking to reflect on variations of a dishes i was preparing for years. It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it

This is why it is so important to hoard offline models. They are already extremely capable, moreso than many realize.

Re: Fable and the end of the free lunch

#236
post #28

What are all these rote coding tasks people do that they can farm it out to lesser models?

"rote" is the wrong framing, the real point is that however sophisticated your task it a lot of it will probably consist of problems they have a good solution already in the training set.

Re: Fable and the end of the free lunch

#237

Earlier quoted context omitted.

It's terrible because it's Llama 3.1 8B. It's such a crappy model because HC1 was a relatively low budget proof of concept. The team that built is working on a better implementation.

Not sure its that to be honest. It seems like maybe its not installed correctly or is like GPT-1/GPT-2 quality? I asked it who is [famous actress] and it started talking about some random person from Mexico with a completely different name. The speed is intoxicating but i'd like for it to actually answer based on what I asked. Thats why I think something might be wrong in implementation on this site. Edit: I went bac…

It has no internet access and only 8B 3-bit parameters. That's simply not enough for it to compress all that much knowledge.

Re: Fable and the end of the free lunch

#239

Earlier quoted context omitted.

What about censorship? > I will be able to use them forever Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?

I dislike all censorship, but US models are much more censored, I often find myself using Chinese models to get answers I want. Now, of course I’d prefer no censoring, but I live in the world we live in. I’m working in the assumption that (like today) there will always be somehow on openrouter, or similar, who will host a model I want to run.

How do Western models censor?

Re: Fable and the end of the free lunch

#240

Earlier quoted context omitted.

I dislike all censorship, but US models are much more censored, I often find myself using Chinese models to get answers I want. Now, of course I’d prefer no censoring, but I live in the world we live in. I’m working in the assumption that (like today) there will always be somehow on openrouter, or similar, who will host a model I want to run.

How do Western models censor?

I’m guessing they’re referring to things like refusals if it thinks your request may be related to building a bioweapon, or hack someone else, etc.
Post reply on HN