Earlier quoted context omitted.
and furthermore, this is because the drivetrain is ~always on the right side of the bike - if you want to inspect or admire a bicycle you look at the right side, as you might look under the hood of a car. (Why the drivetrain is on the right, I don't know. But most bike parts follow open standards so it's quite entrenched.)
I can’t tell you why it’s always on the right , but it’s always on the same side because of network effects. Bicycle frames are not fully symmetric left-right because you need things like a mount point for the derailleur hanger, and optionally affordances to keep the chain off the stays when the wheel is removed. Those things have to be on the same side as the chain. Bikes designed for disc brakes additionally need a…
Muse Spark 1.3
431–440 of 475 posts
Re: Muse Spark 1.3
#432Earlier quoted context omitted.
> Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data. What was the difference between what deepseek did for R1 and what OpenAI did for o1?
openai did human crafted chain of thought dataset training. deepseek didn't have the resources so they attempted RL. doing RL correctly is hard because of the risk of model collapsing.
Re: Muse Spark 1.3
#433Earlier quoted context omitted.
I asked Claude (Opus 4.8) 'If I asked you to "Generate an SVG of a pelican riding a bicycle". What do you think my name would be?' and it immediately knew that this is Simon's go-to benchmark.
I decided to try with each of the options available in Kagi Ultimate, starting with the lower tier models and working my way up until it got it right. Kimi 2.6: treated the question as a riddle, did not know. Kimi 3: Simon Willison GLM 5.3 Flash: "There's no way for me to know that." Going on to say the benchmark is associated with Simon Willison, but I'm more likely to be someone who has just heard of the meme. Clau…
Re: Muse Spark 1.3
#434Earlier quoted context omitted.
By their own benchmarks it is about 10% lower scoring than Qwen 3.6 35b-a3b, but I've added it to my list. Always looking for MoE to compare to it so we can squeeze more out of our local LLM system.
I found it has some "tail" errors, wherein it would make up important details (ie. "happypath.exp" vs "happypaws.exp" and then claim your "DNS is having issues" - where the second domain does not exist), things like that.- ... but correctly supervised it does get some things done.-
Re: Muse Spark 1.3
#435Earlier quoted context omitted.
I found it has some "tail" errors, wherein it would make up important details (ie. "happypath.exp" vs "happypaws.exp" and then claim your "DNS is having issues" - where the second domain does not exist), things like that.- ... but correctly supervised it does get some things done.-
I've found that's generally true of smaller / weaker models. They're quite capable, but you need to distrust them a lot and give them very detailed instructions. Even the free Gemini in Google Search is like this - it lies a lot, clips off important info, and generally goes off the rails if you do too many turns, but it's still very useful if you keep all that in mind.
Re: Muse Spark 1.3
#436Earlier quoted context omitted.
not hard for secrets with explicit patterns and existing pipelines to detect them
Unless they're base64-encoded or compressed?
Re: Muse Spark 1.3
#437Re: Muse Spark 1.3
#438Earlier quoted context omitted.
That's ok to think at this point, given the trajectory of the last few years. Certainly it's one of those things where erring (marginally and slightly) on the side of being safe about it is better than the alternative.
I think where you and I disagree is on whether Anthropic is especially trustworthy on the "safety" front, more trustworthy than various other labs, especially those that produce open models, for example. I simply don't trust Amodei more than I trust, say, Liang Wenfeng. I'm not saying I trust any of them, particularly, I am saying that if a few billionaires have access to this technology, I want access to this techno…
First, Amodei has taken an unusually strong stance among tech companies for not supplying fascist regimes with fascist tooling; in fact, even when threatened with being labeled a national security risk unless he bent the knee, he didn't. Compare and contrast with OpenAI who leapt at the opportunity to bend the knee, or obviously Elon Musk, etc., etc. When you say 'surveillance and control', that's exactly what got Anthropic labeled a national security supply chain risk: Anthropic's unwillingness to be used for that purpose.
Second, it's not clear that giving everyone extremely powerful LLMs is a great idea yet. LLMs can be used for defense and finding vulnerabilities, but that same LLM can be used to create and exploit vulnerabilities, design new lethal weapons, and so on. The history of gun availability in America 'for our freedoms' demonstrates the kind of risk that should be responsibly considered before replicating. And again, there's nuance here; yes, we should not be subjugated by fascist states with sole control of a critical technology obviously; but also, do you trust the median maga 4channer to operate a Mythos-level model with a sense of civilizational responsibility and ethics? It's not an easy and obvious question and it's not as simplistic as your argument would suggest.
Re: Muse Spark 1.3
#439Earlier quoted context omitted.
> and furthermore, this is because the drivetrain is ~always on the right side of the bike While I'm sure this factors into things for advertisements for bike components, there is also just a general preference that westerners have for left-to-right motion. Not just in bike ads, but all ads with (or suggesting) movement. And also not just ads, but movies where directors believe left-to-right motion is associated with…
Research has shown that people like to walk counterclockwise (right to left) through supermarkets, which is why they are arranged like this for maximum profit.
It would be interesting to know how people behave if the entrance is to the left vs right. Would they change the direction they walked though the store, or would they just lose customers due to this "awkward" layout?