Live data from Hacker News

Claude Opus 4.6

anthropic.com

921–930 of 1001 posts

Re: Claude Opus 4.6

#921
post #548

Earlier quoted context omitted.

There was a time when I put the EA-Nasir text into base64 and asked AI to convert it. Remarkably it identified the correct text but pulled the most popular translation of the text than the one I gave it.

Sucks that you got a really shitty response to your prompt. If I were you, the model provider would be receiving my complaint via clay tablet right away.

Imagine you ordered the new Claude Opus and instead you got Gemini telling you to glue the cheese on your pizza...

Re: Claude Opus 4.6

#922
post #90

I'm still not sure I understand Anthropic's general strategy right now. They are doing these broad marketing programs trying to take on ChatGPT for "normies". And yet their bread and butter is still clearly coding. Meanwhile, Claude's general use cases are... fine. For generic research topics, I find that ChatGPT and Gemini run circles around it: in the depth of research, the type of tasks it can handle, and the qual…

Claude itself (outside of code workflows) actually works very well for general purpose chat. I have a few non-technical friends that have moved over from chatgpt after some side-by-side testing and I've yet to see one go back - which is good since claude circa 8 months ago was borderline unusable for anything but coding on the api.

I got my partner using claude for her non technical work. They write a lot of proposals, creates spreadsheets, and occasionally wants some graphs to visualize things. They love that claude creates all of the artifacts right there in the browser and saves them for later in a versioned way.

Re: Claude Opus 4.6

#923

I asked > Can you find an academic article that _looks_ legitimate -- looks like a real journal, by researchers with what look like real academic affiliations, has been cited hundreds or thousands of times -- but is obviously nonsense, e.g. has glaring typos in the abstract, is clearly garbled or nonsensical? It pointed me to a bunch of hoaxes. I clarified: > no, I'm not looking for a hoax, or a deliberate comment on…

When Claude does WebSearch it can delegate it to a sub agent which of it ran in the background will write the entire prompt on a local file and the results. If that happened, I would like to know what it gave you for that. It is always very interesting to know the underlying "recall" of such things. Because often it's garbage in garbage out.

The location might still be on your disk if you can pull up the original Claude JSOn and put it through some `jq` and see what pages it went through to give you and what it did.

Re: Claude Opus 4.6

#924

I asked > Can you find an academic article that _looks_ legitimate -- looks like a real journal, by researchers with what look like real academic affiliations, has been cited hundreds or thousands of times -- but is obviously nonsense, e.g. has glaring typos in the abstract, is clearly garbled or nonsensical? It pointed me to a bunch of hoaxes. I clarified: > no, I'm not looking for a hoax, or a deliberate comment on…

> For my tastes telling me "no" instead of hallucinating an answer is a real breakthrough. It's all anecdata--I'm convinced anecdata is the least bad way to evaluate these models, benchmarks don't work--but this is the behavior I've come to expect from earlier Claude models as well, especially after several back and forth passes where you rejected the initial answers. I don't think it's new.

I can concur that previous models would say "No, that isn't possible" or "No, that doesn't exist". There was one time where I asked it to update a Go module from version X.XX to version X.YY and it would refuse to do so because version X.YY "didn't exist". This back with 3.7 if I recall, and to be clear, that version was released before its knowledge cut off.

I wish I remembered the exact versions involved. I mostly just recall how pissed I was that it was fighting me on changing a single line in my go.mod.

Re: Claude Opus 4.6

#926

I asked > Can you find an academic article that _looks_ legitimate -- looks like a real journal, by researchers with what look like real academic affiliations, has been cited hundreds or thousands of times -- but is obviously nonsense, e.g. has glaring typos in the abstract, is clearly garbled or nonsensical? It pointed me to a bunch of hoaxes. I clarified: > no, I'm not looking for a hoax, or a deliberate comment on…

When Claude does WebSearch it can delegate it to a sub agent which of it ran in the background will write the entire prompt on a local file and the results. If that happened, I would like to know what it gave you for that. It is always very interesting to know the underlying "recall" of such things. Because often it's garbage in garbage out. The location might still be on your disk if you can pull up the original Cla…

this is via the chat, but it did provide this

> The user wants a real, sincere academic paper that looks legitimate — proper journal, real authors with real affiliations, many citations — but is obviously nonsense or garbled. Not a hoax or deliberate prank, but something that was genuinely submitted and published, yet is clearly meaningless if you actually read it. Let me think about what I know and search for examples. There are some famous cases:

> Papers generated by AI/GPT that slipped through - there have been reports of papers with phrases like "Certainly, here is a possible introduction for your topic" or "As an AI language model..." > There have been papers with obvious tortured phrases - "sham" papers that use synonym substitution to avoid plagiarism detection, producing phrases like "sheep pursuing" instead of "shepherd" or "big data" becoming "large information" I'm also thinking of papers that made it through peer review despite being complete gibberish, like the famous SCIgen computer-generated papers that fooled conferences, or papers where the abstract and content don't match at all.

> Now I'm trying to recall specific real examples that fit this criteria — published papers that are genuinely incoherent or nonsensical but appeared in legitimate venues.

This is maybe not the _exhaustive_ search strategy I would have liked, but it's probably something you would do if you were limited to just a certain amount of effort on a task.

Re: Claude Opus 4.6

#927

Earlier quoted context omitted.

> For my tastes telling me "no" instead of hallucinating an answer is a real breakthrough. It's all anecdata--I'm convinced anecdata is the least bad way to evaluate these models, benchmarks don't work--but this is the behavior I've come to expect from earlier Claude models as well, especially after several back and forth passes where you rejected the initial answers. I don't think it's new.

I can concur that previous models would say "No, that isn't possible" or "No, that doesn't exist". There was one time where I asked it to update a Go module from version X.XX to version X.YY and it would refuse to do so because version X.YY "didn't exist". This back with 3.7 if I recall, and to be clear, that version was released before its knowledge cut off. I wish I remembered the exact versions involved. I mostly…

alas, 4.5 often hallucinates academic papers or creates false quotes. I think it's better at knowing that coding answers have deterministic output and being firm there.

Re: Claude Opus 4.6

#928
post #905

Earlier quoted context omitted.

Well, if there are papers that match your criteria, it's hallucinating the "no".

It might be wrong but that’s not really a hallucination. Edit: to give you the benefit of doubt, it probably depends on whether the answer was a definitive “this does not exist” or “I couldn’t find it and it may not exist”

claude said "I want to be straight with you: after extensive searching, I don't think the exact thing you're describing — a single paper that is obviously garbled/badly translated nonsense with no actual content, yet has accumulated hundreds or thousands of citations — exists as a famous, easily linkable example."

Re: Claude Opus 4.6

#930
post #482

Just tested the new Opus 4.6 (1M context) on a fun needle-in-a-haystack challenge: finding every spell in all Harry Potter books. All 7 books come to ~1.75M tokens, so they don't quite fit yet. (At this rate of progress, mid-April should do it ) For now you can fit the first 4 books (~733K tokens). Results: Opus 4.6 found 49 out of 50 officially documented spells across those 4 books. The only miss was "Slugulus Eruc…

[deleted]
Post reply on HN