Viewing profile — conradkay
conradkay
HN member- Joined
- Mon, Dec 31, 2018, 2:42 PM UTC
- HN karma
- 294
- Public activity
- 161 items
- HN profile
- View on Hacker News ↗
About conradkay
No profile information was provided.
Recent public activity
-
comment
Comment #49188208
https://pbs.twimg.com/media/HO-59jQaoAA_JZ1?format=jpg Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price disc…
-
comment
Comment #49187517
Are any of those advantages getting stronger over time? I guess TPUs but Google is selling several gigawatts to Anthropic
-
comment
Comment #49039764
Doing a quick search it seems like the average human score is 49%? I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI …
-
comment
Comment #49039723
I don't think can use the AA index to say something is 10% smarter I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1
-
comment
Comment #49039653
https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-... It seems roughly equal according to Anthropic's benchmarks
-
comment
Comment #49000582
Those are the maximum penalties though It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fract…
-
comment
Comment #49000040
I find it trustworthy since we had Hugging Face's account first: https://huggingface.co/blog/security-incident-july-2026 I don't think they have any real motive to shill OpenAI, pr…
-
comment
Comment #48999758
"Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts…
-
comment
Comment #48999752
Plenty of humans have spent more effort trying to cheat than they would've needed to just do things the right way :)
-
comment
Comment #48999735
https://huggingface.co/blog/security-incident-july-2026 They explain it here, basically for data security/privacy reasons
-
comment
Comment #48969692
Things change fast! For Fable 5 it definitely feels past at least 272k
-
comment
Comment #48969666
It's not quadratic attention, you get that curve from the input tokens going up linearly, since the graph is measuring cumulative cost at each token count. Basically for y=5 it's 5…
-
comment
Comment #48865106
Sol fast isn't the Cerebras 750 tok/s version, it's just 1.5x speed at 2.5x price I assume they didn't use the Cerebras version for this since it's probably very supply-constrained…
-
comment
Comment #48835513
Annoying they didn't show benchmarks for several effort modes, since it seems like it might close the gap with Opus 4.8 by cranking tokens up? Noam Brown (OpenAI) "Implications of …
-
comment
Comment #48835370
I think this one is just a coincidence, bound to happen given the pace of releases For exact timing, probably 10-11am Pacific is just optimal for normal working hours
-
comment
Comment #48739274
Yeah you definitely have to be skeptical regarding sentiment for open/local model capabilities, since there's bias from what people want to be true. I generally agree with this in …
-
comment
Comment #48738853
They should add a Sonnet 5 fast mode at ~Opus pricing
-
comment
Comment #48738452
I think the incentives are less bad since a good chunk of usage comes from subscription plans. There was a fairly major regression in Claude Code performance for some time when the…
-
comment
Comment #48738388
I was surprised to learn that Sonnet generally has the same tokens per second as Opus
-
comment
Comment #48736781
Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable …
-
comment
Comment #48596214
That's for their `JSON` data types. In DuckDB it's just a string meaning lots of queries will have to do JSON parsing on every row, but the inserts are very fast. Definitely a bit …
-
comment
Comment #48576639
It's great but you definitely pay for it. Encoding can be really slow, and to a lesser extent decoding as well. So I still end up using .jpg quite often, or .webp as a good middle …
-
comment
Comment #48575833
My favorite spatial reasoning benchmark: https://minebench.ai/ no tricks, I'd definitely be curious to know how much screenshots help
-
comment
Comment #48561961
They reserved the option to buy it at this price, and are now exercising it
-
comment
Comment #48561785
> If the government takes the bulk of your income after a certain point, there isn't really that big of a push to create ground-breaking technology. I'm skeptical that high taxes i…