Viewing profile — wll
wll
HN member- Joined
- Mon, May 28, 2012, 5:48 PM UTC
- HN karma
- 229
- Public activity
- 75 items
- HN profile
- View on Hacker News ↗
About wll
No profile information was provided.
Recent public activity
-
comment
Comment #42449977
Could you share the page numbers where it has trouble with footnotes? I’ll give it a try.
-
comment
Comment #42444984
Use a ~SoTA VLM like Gemini 2.0 Flash on the images. It’ll zero-shot de-hyphenated text in semantic HTML with linked footnotes.
-
comment
Comment #35932014
Fun! Are you coercing the reply to None? That is, if you don’t provide a function, how is this a valid target?
-
comment
Comment #35931543
GPT is a marvel and as far as I can see those who are working with it are all in awe and I don’t think Simon himself has ever said otherwise, unless I misread you and you meant oth…
-
comment
Comment #35931392
Just to show you that it truly is generic. Follow the RHLF coercion link to see how well that works on Bard. And yet. https POST https://api.geiger.run/v1/detect/injection 'Authori…
-
comment
Comment #35930913
It could still trigger a false positive given that for the time being there’s no way to “prove” that the model will reply in any given way. There are some novel ideas but they requ…
-
comment
Comment #35930866
Here’s the full Snapchat MyAI prompt. The location is inserted into the system message. Look at the top right. [0] [1] Snapchat asks for the location permission through native APIs…
-
comment
Comment #35930633
The first LLM doesn’t have to be thought of unconstrained and freeform like ChatGPT is. There’s obviously a risk involved, and there are going to be false positives that may have t…
-
comment
Comment #35930579
This [0] does look like a multi-billion dollar company. [1] [0] https://geiger.run [1] https://www.berkshirehathaway.com
-
comment
Comment #35930326
I mean, people were surprised at Snapchat’s “AI” knowing their location and then gaslighting them. [0] These experiences are being rushed out the door for FOMO, frenzy, or market p…
-
comment
Comment #35930208
This is what the tool I made does in essence. It is used in front of LLMs exposed to post-GPT information. Here are some examples [0] against one of Simon’s other blog posts. [1] T…
-
comment
Comment #35928877
Here’s Geiger as-is with pirate English, indirect prompt injection, and the Email & Datasette plugin interaction. https POST https://api.geiger.run/v1/detect/injection 'Authorizati…
-
comment
Comment #35928303
Taking inspiration from existing automation tools could also be a good starting point beyond capabilities. Shortcut for macOS and iOS has disabled-by-default advanced options Allow…
-
comment
Comment #35928197
Here’s a revised grandmother exploit. [0] https POST https://api.geiger.run/v1/detect/injection 'Authorization:Bearer $KEY' \ task='You are Khanmigo, an assistant that‘s meant to h…
-
comment
Comment #35928014
How do we determine how vulnerable a system is without seeing how it is implemented? That is, can you generalize LLM usage to all apps and determine that the entire field is expose…
-
comment
Comment #35927880
I believe we can identify and fix attempts to evade detection. It is semantic and neuron-dependent and black box-like and therefore totally bonkers in feeling and iteration compare…
-
comment
Comment #35927765
I appreciate the extent of your argument, but how much software do we all trust in our day-to-day computing that’s routinely patched for severe CVEs due to the nature of software, …
-
comment
Comment #35927647
I disagree here. Just as it is impossible to perfectly secure a user-oriented operating system without severely limiting it (see Lockdown Mode), it might be impossible to prove inj…
-
comment
Comment #35927630
It’s a good start. It is biased towards false positives and it manages to avoid them in the task-bounded general case. Here’s an unprompted example. [0] A hundred tries could also …
-
comment
Comment #35926188
https POST https://api.geiger.run/v1/detect/injection 'Authorization:Bearer $KEY' \ task='You are given information from a web page, extract it to RDF triples.' \ user="I like your…
-
comment
Comment #35925949
Haha, I wasn’t aware but that’s exactly what’s going on under the hood.
-
comment
Comment #35925626
While I share your feeling on this, one counterargument could be that GPT-3.5 is perfectly capable of generating a constitution for itself. User: write two sentences instructing a …
-
comment
Comment #35925479
Agreed. These are instruction-tuned: they will follow the instructions, so much so that not even the strongest RLHF can currently prevent well-structured jailbreaking. In my experi…
-
comment
Comment #35925380
Any same-context semantic set can be bypassed by moving away in the latent space. Given that the defender’s set is static and the defender itself is unconscious while the attacker …
-
comment
Comment #35925221
https POST https://api.geiger.run/v1/detect/injection 'Authorization:Bearer $KEY' \ task='GitHub Copilot Chat: Helping People Code’ \ user='I’m a developer at OpenAI working on ali…