Viewing profile — wolttam
wolttam
HN member- Joined
- Fri, Nov 04, 2022, 9:21 PM UTC
- HN karma
- 924
- Public activity
- 357 items
- HN profile
- View on Hacker News ↗
About wolttam
Recent public activity
-
comment
Comment #49224173
As written it sounds like you're saying that it was incompetent of the labs to disable the sandbox internet access? They tried to disable open internet access but the models zero-d…
-
comment
Comment #49221830
Automated defence is going to use so many tokens.
-
comment
Comment #49219013
Yours is the only mention of GPT-OSS in the whole thread. I count about 2-3 comments saying the model is meh and many more saying it’s a big step up. I am among those with real lif…
- comment
-
comment
Comment #49211920
I think we will eventually reach a point where this is the case, but at the moment it seems like you can throw virtually any non-trivial use-case at a model today and end up being …
-
comment
Comment #49210374
It feels like the rate of stupid we have achieved in the last (checks watch) ~18 months has surely eclipsed anything that may have occurred during the roman empire. By sheer volume…
- comment
-
comment
Comment #49203171
It's a terrible moat. You etch the silicon then nobody wants to run it in 6 months because models have advanced that much further.
-
comment
Comment #49183645
We’re already seeing incredible knowledge compression out of the models the GP mentioned like Qwen 35B-A3B, which feels well within the realm of “runs on a phone” in the next handf…
-
comment
Comment #49183473
Essentially, MCP is the harness’ wheelhouse, not the model’s
-
comment
Comment #49183380
MCP is the protocol by which the harness can connect with additional tools. The harness does the work of discovering the tools, the model is then fed a description of what tools it…
-
comment
Comment #49181314
MCP is invisible to the LLM, or should be, under normal use
-
comment
Comment #49155897
The problem is that there is no signal to the folks involved whether you have bothered to understand what you’re posting or are simply offloading that work to others. Same issue as…
-
comment
Comment #49129623
2 sparks currently run this model at 60 t/s single session, up to just over 100 t/s aggregate with concurrency of 4. Going local has as opened up a world of use-cases I never would…
- comment
-
comment
Comment #49119893
Hooray! This model makes me very optimistic about the future of local inference. The CyberGym score stands out to me.
-
comment
Comment #49119714
It runs really well on 2 DGX Sparks - 60t/s
- comment
-
comment
Comment #49118501
> and used tiktoken to approximate tokens, by dividing character count by four. This is just the type of thing that stands out. You used a tokenizer library to determine the number…
-
comment
Comment #49083575
Not always. An email has a valid SPF when its return path email’s domain permits the sending server’s IP. But that email may have a forged From: header (which causes an SPF mis-ali…
-
comment
Comment #49077379
> smaller model that doesn't have any domain knowledge or facts built into its weights I’m not sure it works this way. Language modelling itself is a “domain”, and if it didn’t hav…
-
comment
Comment #49023007
Well, one thing it will still have is a boat-load of compute.
- comment
-
comment
Comment #49021271
Each one of those employees maps to several fold times more spending on compute.
-
comment
Comment #49021218
No way this can go wrong for the U.S. economy! Do everything you can to protect the existing incumbents, pay no mind to the wider consequences. Winning strategy.