Live data from Hacker News

Viewing profile — usaar333

usaar333

HN member
Joined
Thu, Apr 03, 2008, 1:47 AM UTC
HN karma
4,075
Public activity
1,815 items

About usaar333

No profile information was provided.

Recent public activity

  1. comment
    Comment #49053696

    > and eventually AI model will be commodified This axiom not being true (and I'd bet against it) means your overall conclusion is false.

  2. comment
    Comment #48231831

    Tech workers get paid in equity and many in the semiconductor industry are making far far more than this a year with all the equity appreciation.

  3. comment
    Comment #48149944

    How does delaying the release not solve anything? It puts everyone on a notice to fix all security vulnerabilities now

  4. comment
    Comment #48016456

    I hate waiting on hold for 30 minutes even more.

  5. comment
    Comment #48016455

    There's literally a link on the blog post to an article noting they hit $150M ARR.

  6. comment
    Comment #48016445

    Voice agents have capabilities and policy to alter customer state. Just the other day I called into a CC company and the AI waived an interest charge.

  7. comment
    Comment #47795453

    page is updated to state: MCP-Atlas: The Opus 4.6 score has been updated to reflect revised grading methodology from Scale AI.

  8. comment
    Comment #47736190

    > But even setting aside the leaked answers, the scorer’s normalize_str function strips ALL whitespace, ALL punctuation, and lowercases everything before comparison. This means: I …

  9. comment
    Comment #46956327

    True, but it gets you higher accuracy. Gemini had the best aa-omniscience score https://artificialanalysis.ai/evaluations/omniscience

  10. comment
    Comment #46902793

    Openai has; they don't even mention score on gpt-5.3-codex. On the other hand, it is their own verified benchmark, which is telling.

  11. comment
    Comment #46902501

    i'd interpret that as rounding error. that is unchanged swe-bench seems really hard once you are above 80%

  12. comment
    Comment #46016627

    In Quebec it was a 20% jump in mother employment: https://www.bloomberg.com/news/articles/2018-12-31/affordabl... And had all sorts of negative outcomes for the kids: https://www.e…

  13. comment
    Comment #45969455

    claude 4.5 gets 82% on their own highly customized scaffolding. (parallel compute with a scoring function). That beats Doubao

  14. comment
    Comment #45454889

    That wasn't a ceasefire violation. It was a six week ceasefire that had expired at the beginning of March

  15. comment
    Comment #45430267

    Physics seems better than veo 3 at least from demo videos

  16. comment
    Comment #45417855

    Except it is sublinear. Sonnet 4 was 10.2% above sonnet 3.7 after 3 months.

  17. comment
    Comment #44831916

    No it doesn't. If it were even linear compared to o1 -> o3, we'd be at 2.43 hours. Instead we're only at 2.29. Exponential would be at 3.6 hours

  18. comment
    Comment #44828550

    No, this is below expectations on both Manifold and lesswrong ( https://www.lesswrong.com/posts/FG54euEAesRkSZuJN/ryan_green... ). Median was ~2.75 hours on both (which already rep…

  19. comment
    Comment #44827226

    At this point the prediction for SWE bench (85% by end of this month) is not materializing. We're actually quite far away.

  20. comment
    Comment #44800664

    No obvious gains I feel from quick chats, but too early to tell. These benchmark gains aren't that high, so I doubt it is that obvious.

  21. comment
    Comment #44790197

    > Firstly, if your prior is that every previous startup failed, what does that say about your future chances of success? The prior is the market. It isn't sane to use your own prio…

  22. comment
    Comment #44786076

    Why is modal return so important? You'll work more than 2 jobs

  23. comment
    Comment #44781598

    It's a probabilistic model. It assumes (correctly) that the low probability of a home run times the home run's valuation is quite large ("expected returns" in the probabilistic sen…

  24. comment
    Comment #44781307

    The value of the equity package is 4x higher than the FAANG equivalent equity package (at preferred/market pricing) - that's not the same as saying the shares themselves are worth …

  25. story