Live data from Hacker News

Viewing profile — timbilt

timbilt

HN member
Joined
Fri, Nov 01, 2024, 8:50 AM UTC
HN karma
571
Public activity
53 items

About timbilt

No profile information was provided.

Recent public activity

  1. story
  2. story
    Google open-sources experimental agent orchestration testbed Scion

    https://googlecloudplatform.github.io/scion/overview/

  3. story
  4. story
  5. comment
    Comment #44834159

    Yes, but in a case like this it's a neutral third-party running the benchmark. So there isn't a direct incentive for them to favor one lab over another. With public benchmarks we'r…

  6. comment
    Comment #44834088

    > Unlike many public benchmarks, the PR Benchmark is private, and its data is not publicly released. This ensures models haven’t seen it during training, making results fairer and …

  7. story
  8. story
  9. story
  10. story
  11. story
  12. story
  13. story
  14. comment
    Comment #43242945

    anyone else concerned that training models on synthetic, LLM-generated data might push us into a linguistic feedback loop? relying on LLM text for training could bias the next mode…

  15. story
  16. story
  17. comment
    Comment #43004425

    Twitter thread about this by the author: https://x.com/jonasgeiping/status/1888985929727037514

  18. story
  19. story
  20. story
  21. story
  22. comment
    Comment #42907357

    Until we get real-time learning to work in production, every AI tool feels like it's getting dumber over time. It goes very quick from "wow this is magic" to starting to notice all…

  23. comment
    Comment #42870769

    The weirdness of LLMs is that they're so damn good at so many things but then you see these glaring gaps that instantly make them seem dumb. We desperately need benchmarks and eval…

  24. story
  25. story