Live data from Hacker News

Viewing profile — conradkay

conradkay

HN member
Joined
Mon, Dec 31, 2018, 2:42 PM UTC
HN karma
294
Public activity
161 items

About conradkay

No profile information was provided.

Recent public activity

  1. comment
    Comment #49188208

    https://pbs.twimg.com/media/HO-59jQaoAA_JZ1?format=jpg Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price disc…

  2. comment
    Comment #49187517

    Are any of those advantages getting stronger over time? I guess TPUs but Google is selling several gigawatts to Anthropic

  3. comment
    Comment #49039764

    Doing a quick search it seems like the average human score is 49%? I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI …

  4. comment
    Comment #49039723

    I don't think can use the AA index to say something is 10% smarter I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1

  5. comment
    Comment #49039653

    https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-... It seems roughly equal according to Anthropic's benchmarks

  6. comment
    Comment #49000582

    Those are the maximum penalties though It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fract…

  7. comment
    Comment #49000040

    I find it trustworthy since we had Hugging Face's account first: https://huggingface.co/blog/security-incident-july-2026 I don't think they have any real motive to shill OpenAI, pr…

  8. comment
    Comment #48999758

    "Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts…

  9. comment
    Comment #48999752

    Plenty of humans have spent more effort trying to cheat than they would've needed to just do things the right way :)

  10. comment
    Comment #48999735

    https://huggingface.co/blog/security-incident-july-2026 They explain it here, basically for data security/privacy reasons

  11. comment
    Comment #48969692

    Things change fast! For Fable 5 it definitely feels past at least 272k

  12. comment
    Comment #48969666

    It's not quadratic attention, you get that curve from the input tokens going up linearly, since the graph is measuring cumulative cost at each token count. Basically for y=5 it's 5…

  13. comment
    Comment #48865106

    Sol fast isn't the Cerebras 750 tok/s version, it's just 1.5x speed at 2.5x price I assume they didn't use the Cerebras version for this since it's probably very supply-constrained…

  14. comment
    Comment #48835513

    Annoying they didn't show benchmarks for several effort modes, since it seems like it might close the gap with Opus 4.8 by cranking tokens up? Noam Brown (OpenAI) "Implications of …

  15. comment
    Comment #48835370

    I think this one is just a coincidence, bound to happen given the pace of releases For exact timing, probably 10-11am Pacific is just optimal for normal working hours

  16. comment
    Comment #48739274

    Yeah you definitely have to be skeptical regarding sentiment for open/local model capabilities, since there's bias from what people want to be true. I generally agree with this in …

  17. comment
    Comment #48738853

    They should add a Sonnet 5 fast mode at ~Opus pricing

  18. comment
    Comment #48738452

    I think the incentives are less bad since a good chunk of usage comes from subscription plans. There was a fairly major regression in Claude Code performance for some time when the…

  19. comment
    Comment #48738388

    I was surprised to learn that Sonnet generally has the same tokens per second as Opus

  20. comment
    Comment #48736781

    Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable …

  21. comment
    Comment #48596214

    That's for their `JSON` data types. In DuckDB it's just a string meaning lots of queries will have to do JSON parsing on every row, but the inserts are very fast. Definitely a bit …

  22. comment
    Comment #48576639

    It's great but you definitely pay for it. Encoding can be really slow, and to a lesser extent decoding as well. So I still end up using .jpg quite often, or .webp as a good middle …

  23. comment
    Comment #48575833

    My favorite spatial reasoning benchmark: https://minebench.ai/ no tricks, I'd definitely be curious to know how much screenshots help

  24. comment
    Comment #48561961

    They reserved the option to buy it at this price, and are now exercising it

  25. comment
    Comment #48561785

    > If the government takes the bulk of your income after a certain point, there isn't really that big of a push to create ground-breaking technology. I'm skeptical that high taxes i…