Viewing profile — espadrine
espadrine
HN member- Joined
- Mon, Sep 12, 2011, 7:47 AM UTC
- HN karma
- 7,476
- Public activity
- 1,285 items
- HN profile
- View on Hacker News ↗
About espadrine
Recent public activity
-
comment
Comment #49241156
Models start going to extreme, damaging lengths to achieve ambiguous prompts[0]. Having good sandboxes is now a must IMO. But sbx is a bit annoying to use with OpenCode for instanc…
-
comment
Comment #49113412
I maintain this meta-benchmark leaderboard: https://metabench.organisons.com/ With this new price change, Terra does look pretty Pareto’ed by Luna. On agentic coding, pairing Sol M…
-
comment
Comment #48843941
Composer 2.5 finetuned Kimi K2.5[0]. In the blog post, it is unclear whether Grok 4.5 is also a finetune on top of Kimi; they do imply it is also a finetune. > Training included tr…
-
comment
Comment #48720320
Iridum gains 23 launches per year with 100% success rate in the past 12 months, a satellite manufacturing pipeline with 6 satellites produced and launched, and a cost-to-orbit of $…
-
comment
Comment #48426197
I see it mean two things: 1. Indeed, Google is compute-constrained, and is ready to buy any it can. 2. xAI (now SpaceXAI) has a lot of idle compute, which it resells to Cursor, Ant…
-
comment
Comment #48268038
I use Le Chat as default search engine, using this search engine string: https://chat.mistral.ai/chat?q=Give%20a%20list%20of%20links%... (In most browsers, you can input any URL wi…
-
comment
Comment #48195217
His goal could simply be to learn SOTA architectures. When rumors started that GPT-4 design would be kept secret, he likely wanted to know what architecture it would be. Perhaps he…
-
comment
Comment #48170377
[dead]
-
comment
Comment #48169572
Could you link to a project that you consider the best Tailwind use you know? I have a bias against Tailwind, admittedly because I saw some vibecoded Tailwind where each class was …
-
comment
Comment #48010478
> A portable battery should be considered to be removable by the end-user when it can be removed with the use of commercially available tools and without requiring the use of speci…
-
comment
Comment #47932560
How much did this pretraining run cost? I am impressed that it is now practical to do such efforts. Let me try a guess for the cost; please fact-check it if you can. They indicate …
-
comment
Comment #47890752
I have a rebuttal to your rebuttal. Models somehow have a shared identity. Pretraining causes them to generate “AI chatbot” as a concept, and finetuning causes them to identify wit…
-
comment
Comment #47410809
Input: Following overhiring during COVID, we are laying off workers but claim it is because of AI. As we continue to evolve in this rapidly shifting landscape, we are making the di…
-
comment
Comment #47150118
Interestingly, while it uses diffusion, it generates incorrect information, and it doesn't fix it when later in the text it realizes that it is incorrect: > The snail you’re likely…
-
comment
Comment #47085749
AI companies have two conflicting interests: 1. curating the default personality of the bot, to ensure it acts responsively; 2. letting it roleplay, which is not just for the paras…
-
comment
Comment #46902378
It is quite impressive. I have seen the same impressive performance about 7 months ago here: https://kyutai.org/stt If I look at the architecture of Voxtral 2, it seems to take a p…
-
comment
Comment #46600299
Does Apple develop a competing search engine?
-
comment
Comment #46598812
Counterpoint: iOS’s biggest competitor is Android. They are now effectively funding their competition on a core product interface. I see this as strategically devastating.
-
comment
Comment #46568923
My bar for super-rough is Servo, which doesn't have password autofill… and doesn't render the Orion page right. Orion is less rough, but the color scheme doesn't work, and it doesn…
-
comment
Comment #46113600
Good question. There's 2 points to consider. • For both Kimi K2 and for Sonnet, there's a non-thinking and a thinking version. Sonnet 4.5 Thinking is better than Kimi K2 non-thinki…
-
comment
Comment #46111351
Two aspects to consider: 1. Chinese models typically focus on text. US and EU models also bear the cross of handling image, often voice and video. Supporting all those is additiona…
-
comment
Comment #46079853
Indeed. A mouse that runs through a maze may be right to say that it is constantly hitting a wall, yet it makes constant progress. An example is citing Mr Sutskever's interview thi…
-
comment
Comment #45659976
That makes sense. Why RVQ though, rather than using the raw VAE embedding? If I compare rvq-without-quantization-v4.png with rvq-2-level-v4.png, the quality seems oddly similar, bu…
-
comment
Comment #45484711
> DeepSeek models cost more to use than comparable U.S. models They compare DeepSeek v3.1 to GPT-5 mini. Those have very different sizes, which makes it a weird choice. I would exp…
-
comment
Comment #45415667
Input: $0.07 (cached), $0.56 (cache miss) Output: $1.68 per million tokens. https://api-docs.deepseek.com/news/news250929