Viewing profile — timbilt
timbilt
HN member- Joined
- Fri, Nov 01, 2024, 8:50 AM UTC
- HN karma
- 571
- Public activity
- 53 items
- HN profile
- View on Hacker News ↗
About timbilt
No profile information was provided.
Recent public activity
- story
-
story
Google open-sources experimental agent orchestration testbed Scion
https://googlecloudplatform.github.io/scion/overview/
- story
- story
-
comment
Comment #44834159
Yes, but in a case like this it's a neutral third-party running the benchmark. So there isn't a direct incentive for them to favor one lab over another. With public benchmarks we'r…
-
comment
Comment #44834088
> Unlike many public benchmarks, the PR Benchmark is private, and its data is not publicly released. This ensures models haven’t seen it during training, making results fairer and …
- story
- story
- story
- story
- story
- story
- story
-
comment
Comment #43242945
anyone else concerned that training models on synthetic, LLM-generated data might push us into a linguistic feedback loop? relying on LLM text for training could bias the next mode…
- story
- story
-
comment
Comment #43004425
Twitter thread about this by the author: https://x.com/jonasgeiping/status/1888985929727037514
- story
- story
- story
- story
-
comment
Comment #42907357
Until we get real-time learning to work in production, every AI tool feels like it's getting dumber over time. It goes very quick from "wow this is magic" to starting to notice all…
-
comment
Comment #42870769
The weirdness of LLMs is that they're so damn good at so many things but then you see these glaring gaps that instantly make them seem dumb. We desperately need benchmarks and eval…
- story
- story