Live data from Hacker News

Anthropic surpasses OpenAI to become most valuable AI startup

qazinform.com

491–500 of 512 posts

Re: Anthropic surpasses OpenAI to become most valuable AI startup

#491
post #437
post #420

Earlier quoted context omitted.

> I just don't believe non-deterministic tools can actually be benchmarked. It's all hoopla to me. We benchmark non-deterministic things all the time and it's frankly not even that unusual or hard. You yourself indicate that one model outperforms another one in your experience on various facets, and that is itself a benchmark. The more relevant question is probably how well does a given benchmark translate to improve…

My anecdotal experience isn't a benchmark. Just because I feel like something is better or different doesn't mean it actually is.

> Just because I feel like something is better or different doesn't mean it actually is.

Of course, but it is a data point, and multiple such data points can be aggregated. This is true even if all you can do is compare two things.

The shape of that data will reveal something more about the thing you want to measure than the null hypothesis you'd otherwise have.

Re: Anthropic surpasses OpenAI to become most valuable AI startup

#494

Ah, it’s a good time to check in with gwern on our conversation about oAI vs Anthropic: https://news.ycombinator.com/item?id=40816755 and our predictions (ca two years ago). Upshot - poetry expertise does not seem to be the primary focus these days, perhaps to the detriment of the entire world. We did move on from training scaling to “test time” scaling (which I hate as a name btw), Ilya does not seem to have been ne…

Mechinterp in general is just completely undervalued right now (and agreed Anthropic's team is doing the most rigorous work, now accompanied by Goodfire). They're doing the closest work to neuroscience's in vivo 'thought-tracing', which is just the most wild science fiction sort of thing to be working on, and yet I feel the average person has no idea this sort of work is happening. When combined with the idea of the…

it's not undervalues, many people are working on it following anthropic's lead. It just doesn't seem to be any useful, so it's even overvalued

Re: Anthropic surpasses OpenAI to become most valuable AI startup

#495

I never want to hear from developers again that they are not susceptible to marketing. I see meet ups specifically about Claude often. Modern tupperware party. A colleague was convinced Claude is better so we played a game. We used the claude code and codex harness and I implemented some prs they needed with gpt5.5 and opus4.7 and asked them to identify which came from which only from the code. Couldn’t tell. Edit: i…

I don't think the success is due to marketing. They've been top of the LLM Arena leaderboard for most of the year which I think is blind AB testing. Most people on HN say they are best for code. I've never seen their marketing. Your post was the first time I even realised it exists.

Re: Anthropic surpasses OpenAI to become most valuable AI startup

#496
post #104

I think Sam Altman is an asshole and I prefer to spend my money elsewhere. Frontier models being commoditize is inevitable. OpenAI thinks they're still competing on technology, and not user experience and market reputation otherwise they'd understand the continuous negative PR generated by Altman's chaos is going to cost them everything.

How can you say this as if supporting Dario is any better. At the top level of anything there is almost no such thing as a non-asshole. None of them care genuinely about you they just want your money.

Hassabis seems decent. People have strange attitudes to tech leaders like they are all superhuman or evil or all aresholes. They are just normal people who got the top job, some good some bad, some mediocre.

Re: Anthropic surpasses OpenAI to become most valuable AI startup

#498
post #451

Earlier quoted context omitted.

The cost/performance is terrible for higher end cards. In a few years your card is now worth nothing because the lower end cards of the next gen are matching it and there's yet another new SOTA card out. So people end up buying that and chasing the dragon. The 4k situation is a good point because nvidia deliberately don't provide 24GB except the 90 series, but ... you're too good for DLSS? You can't move a texture sl…

Someone who bought a 4090 "a few years" ago can now sell it for more than they paid for it, but never mind that.

Do you think GPUs are going to keep going up in value forever like houses? The current situation (game console price rises years into their lifecycle) is unprecedented and irrelevant to my point.

And if you need SOTA then you can sell your old card sure, but the next xx90 card is now 2x the price as the last gen. So you're not any better off.

Re: Anthropic surpasses OpenAI to become most valuable AI startup

#499
post #353
post #26

Earlier quoted context omitted.

This is like saying you gave a Taylor Swift fan sheet music from 1984 and from Michael Jackson’s thriller and they couldn’t tell the difference. I have a strong affinity for Claude Code because of the interaction experience and overall tone / vibe / process. I am 100% willing to believe the code it produces is identical or possibly less good than Codex. I enjoy working with Claude in a way I just don’t get from OpenA…

Can you give me some examples of these interactions / vibe?

I’m working on an AI-assisted music composition and criticism tool (giant project, may or may not pan out). It covers audio (samples, levels, etc), theory (harmony, rhythm), genre (classical, Motown), song intent, etc.

Working on the melody model, I asked Claude to thread it through those dimensions, both for composition and analysis. It’s a tough problem because there are heuristics but not rules for melody, so you have to come at it in layers: pitch and dynamics for analysis, intent -> genre -> harmony for composition.

Lots of research and brainstorming, and I like that Claude will start implementing a plan we decided on, then say “hey wait this isn’t making sense” and pivot or change scope in sensible ways. For instance (btw its recent obsession with “honesty” is driving me batty, so that’s a good counter example right there):

> The honest fix: distinguishing shaped-from-aimless is genuinely profile-relative (it needs the declared intent) — so v1 should not verdict it. v1 classifies only what’s genre-safe to call without intent (static vs active), reports all the facts that feed the shaped/random judgment (step/leap, reversal, contour, alphabet), and defers the shaped-vs-random verdict to profile-relative phase 2b. This is the same honesty as performance deferring declared-profile grading to 2b — and it’s more deferred for melody because melody is more genre-relative than feel. Let me revise the lens accordingly.

Re: Anthropic surpasses OpenAI to become most valuable AI startup

#500

Earlier quoted context omitted.

Haven't heard about the universal subspace hypothesis yet, so I appreciate the digression.

Ya, super interesting research area the authors explored of basically trying to answer the question: "Is there a canonical/intrinsic way that concepts/representations/information are 'stored' in the universe/reality?". They tested that by performing "spectral analysis of over 1100 models - including 500 Mistral-7B LoRAs, 500 Vision Transformers, and 50 LLaMA-8B models ... by applying spectral decomposition techniques…

[dead]
Post reply on HN