Viewing profile — saladtoes
saladtoes
HN member- Joined
- Tue, Jan 29, 2019, 3:10 PM UTC
- HN karma
- 5
- Public activity
- 23 items
- HN profile
- View on Hacker News ↗
About saladtoes
No profile information was provided.
Recent public activity
- story
-
comment
Comment #44086218
Agreed on CaMeL as a promising direction forward. Guardrails may not get 100% of the way but are key for defense in depth, even approached like CaMeL currently fall short for text …
-
comment
Comment #44086178
https://www.lakera.ai/blog/claude-4-sonnet-a-new-standard-fo... These LLMs still fall short on a bunch of pretty simple tasks. Attackers can get Claude 4 to deny legitimate request…
- story
- story
- story
-
comment
Comment #35944421
I've been playing Gandalf in the last few days, it does a great job at giving an intuition for some of the subtleties of prompt engineering: https://gandalf.lakera.ai Thanks for pu…
- story
- story
- story
- story
-
comment
Comment #33436960
Is that really a universal fact? In any case, my statement goes in the opposite direction: is a reliable system necessarily "well understood", in the sense that it can explain its …
-
comment
Comment #33435239
Explainability is not a given in many more traditional complex systems. Decisions are often an aggregation of a large number of signals, and one can often not conceive of a single …
-
comment
Comment #33434360
Interesting, thanks for sharing! Somehow I'm not surprised. My experience building systems for real world applications is that choosing a simple CNN is usually the way to go, and i…
- story
- story
- story
- story
-
comment
Comment #27747836
“The latter being when the training or test data follows a different distribution to the in-operation data” This form of ML bug is the most challenging to catch. The true in-operat…
- story
- story
-
comment
Comment #26831692
Spot on, the EU seems to agree: https://techcrunch-com.cdn.ampproject.org/c/s/techcrunch.com...
-
comment
Comment #26749020
Oh no!! I really wanted to know how they develop their AI so reliably :). This is a massive issue today. Hope we develop tools and processes to get us there soon.”