Viewing profile — 6thbit
6thbit
HN member- Joined
- Wed, Nov 04, 2020, 12:04 AM UTC
- HN karma
- 746
- Public activity
- 229 items
- HN profile
- View on Hacker News ↗
About 6thbit
No profile information was provided.
Recent public activity
-
comment
Comment #49200073
So what Meta believes fair for "paying" for your data is $0.1/Mtok plus the opportunity cost of $3/Mtok in output?
-
comment
Comment #49200008
is this their way to push their harness? people interested in the discounted -contributor model would increase their harness beta testers. But it'd be a terrible strategy to captur…
-
comment
Comment #49199916
This is likely an legitimate risk, not from you exactly, but if there's unethical competitors with unlimited budgets, their training could indeed be poisoned. Not sure if anyone wo…
-
comment
Comment #49162770
So we could run a lighter LLM in front of humans, which translates from 'no domain knowledge' to 'domain expert' and in turn prompts over to the larger LLM. Then the larger LLM get…
-
comment
Comment #49159027
Use it not just as an output tool but as an input as well. You could try doing the high level design yourself at least. Ask for its review and iterate without asking it to do it al…
-
comment
Comment #49128123
I've always found tailscale's json config a bit intimidating. I greatly appreciate the new UI that makes it easier to define rules, alas, both going to relevant docs straight from …
-
comment
Comment #49117490
There was no rush for this disclosure on their side. And they publish at a point where they have not yet taken corrective actions: > Some of the solutions here may even be simple f…
-
comment
Comment #49117427
> closer to a harness and operational failure than a model alignment failure. Our models were told they had no internet access and to capture the flag, while in fact being misconfi…
-
comment
Comment #49090965
How does it fare against its own codebase?
-
comment
Comment #49040724
Honestly that's the simplest explanation and thus likely the correct one.
-
comment
Comment #49039135
Anyone has an insight into how much money labs are putting into benchmarks? Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!). Likely just a drop in t…
-
comment
Comment #49039084
"although Opus 5 shows improvements in its ability to identify software vulnerabilities, it is substantially behind Mythos 5 in its ability to exploit them." "Opus 5’s safeguards m…
-
comment
Comment #49039052
Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most_…
-
comment
Comment #49028348
[dead]
-
comment
Comment #49014543
I think svg is a balanced test because of the level of indirection and the required 'conceptualization' of physical elements then expressed through code.
-
comment
Comment #49014541
What would be an alternative format or process with a similar effort to drawing SVGs?
-
comment
Comment #49011286
Presumably this was Sol on xhigh, then over to Pro (as per his indication on chat)? Is there any way to tell a conversation's model and thinking level?
-
comment
Comment #49010358
and that scientific evidence is new? like a new study boosted this or just resurfacing?
-
comment
Comment #49009041
Why is creatine so popular lately? what changed?
-
comment
Comment #49002223
Did a human ask it to abuse vulnerabilities and escalate across external systems?
-
comment
Comment #49001034
The agent did it intentionally and willfully and knowingly. But you can’t sue the agent, I suppose. And the human didn’t ask the agent to do so.. so not a problem? Or the legislati…
-
comment
Comment #49000991
If an individual did this, a massive CFAA hammer would be falling on their heads. Even though it doesn’t seem to be the case, OpenAI could’ve been trying to hack into HF and blame …
-
comment
Comment #48928804
I feel this as more of a fashion runway garment type object, more statement and vision than what you'd see in everyday retail. Reminds me of an italian specialty coffee shop that p…
-
comment
Comment #48836203
Likely doesn’t make sense, at least not immediate/mid term. They don’t have to aim for number one though, just for enough cash flow and growth.
-
comment
Comment #48835045
Can it delegate to just one agent at a time or can it spawn multiple subagents for different tasks?