Live data from Hacker News

Viewing profile — osti

osti

HN member
Joined
Tue, Jan 05, 2016, 8:42 PM UTC
HN karma
574
Public activity
235 items

About osti

No profile information was provided.

Recent public activity

  1. comment
    Comment #49158271

    It's an axiom that the modern Western mind is built upon.

  2. comment
    Comment #49127389

    Do you have any numbers on the solve quality? Exploitability numbers etc.

  3. comment
    Comment #49026386

    You can use geekbench 5 in that case. But given that they deprecated that, it might be harder to compare to others.

  4. comment
    Comment #49026366

    I agree. For me personally I mostly only care about single thread geekbench variant, I believe it's an excellent proxy for general performance of a CPU. Multi thread geekbench (or …

  5. comment
    Comment #48994037

    In computer science, one definition of algorithm is basically any program that runs on a turing machine. By that definition, any LLM is an algorithm.

  6. comment
    Comment #48980681

    Lol yet I've used Apple and Android phones extensively and would choose Android every single time.

  7. comment
    Comment #48963492

    He's talking about the plans, you are talking about API prices.

  8. comment
    Comment #48963272

    No idea lol, didn't even know those exist..

  9. comment
    Comment #48963250

    It is complicated, but paying for the cheaper usd plans really don't get you much usage.

  10. comment
    Comment #48963242

    Nah that won't work. I don't know tbh, I just used someone else's number.

  11. comment
    Comment #48961996

    Absolutely do not pay for the kimi plans thinking they will be cheaper. If you sign up with a Chinese phone number, you can get the same plan for 200 yuan instead of 200 usd, it al…

  12. comment
    Comment #48961930

    GPT should be better at these optimization problems given that they won the recent atcoder heuristics competition against top humans. And Anthropic is less focused on these types o…

  13. comment
    Comment #48851663

    This was extremely impressive to me. AtCoder has the hardest problems these days, usually the human onsite final round contestants can't solve more than 2 or 3 problems. This year …

  14. comment
  15. comment
    Comment #48849903

    Mythos probably wouldn't, otherwise they'd have included it in their release. Next version of Mythos probably will though. And yeah.. Reality has not been kind to LeCun.

  16. comment
    Comment #48849777

    SWE-bench series just aren't that great by today's standard, even Anthropic previously stated Claude had memorized solutions for the non Pro version of the benchmark, I suspect the…

  17. comment
    Comment #48849553

    Huh so that's why it's hard to find. They probably haven't properly optimized their caching, or they are just trying to make more money from there.

  18. comment
    Comment #48849386

    And they'd be right, it's an almost saturated benchmark where even some subpar open source models score very well on. And most models are clustered within a small range so it reall…

  19. comment
    Comment #48849363

    SWE-Bench pro is pretty much useless now even though many ppl still look at it. OpenAI published a report yesterday saying so as well. Only look at DeepSWE and FrontierCode right n…

  20. comment
    Comment #48849251

    GPT usually performs better on DeepSWE while Claude does better on FrontierCode. These two coding benchmarks are pretty much the only ones right now that's still worth taking a loo…

  21. comment
    Comment #48813643

    Most programming is that, but most music and literature are probably uninspired junks as well. But there are many beautiful algorithms (such as the ones in Knuths books) that are m…

  22. comment
    Comment #48752700

    I think at the current stage of LLM, it just doesn't make sense to ever have an annual sub. Things change too quickly that you really don't want to be stuck with one model.

  23. comment
    Comment #48741187

    If anything that'll be more obnoxious because they have to show the government that it's safe.

  24. comment
    Comment #48721988

    This announcement also mentioned that they will release the next version (official non preview version) of v4 in mid July.

  25. comment
    Comment #48714062

    Is it only the Indian government? I don't think that's in any way unique to India, I've seen many poor government websites.