Earlier quoted context omitted.
>They're so, so incomprehensible because the LLM has a super limited theory of mind for readers. They always assume that external readers have access to the full context and history of decisions in the project development The transformer does not yet understand the non-transformer.[0] This is probably because all the data we trained it on was created by non-transformers, so it thinks it's a non-transformer, but it is…
I think the models would need far far more introspection for the problem to be it understanding how it thinks but not how others think. I really doubt it understands how it thinks.
Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
151–160 of 181 posts
Re: Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
#152Re: Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
#153Earlier quoted context omitted.
Rewrite the readme in my voice, use all of my comments and responses in current context as source. Don't use em dashes or other LLM things. Done, solved. Never had a problem with a readme or email since.
The emdash allergy is shortcut sheep mindset. Emdashes are commonly accepted and orthographically correct punctuation in many languages. Also it looks like English might not be the author’s first language—might not lead to the best training corpus.
Re: Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
#154Earlier quoted context omitted.
i thought, here on HN, we were past the "oooh it's written by a LLM it's bad!". I care about the craft, well designed systems, good clean architecture and code, etc... But i also care about reaching goals. Whether i do it working on my own, or with human coworkers or with AI coworkers doesn't matter that much to me. Yes, the result is sometimes the most important thing.
> I care about the craft No you don't. Buying a table and sanding the edges off doesn't mean you're a carpenter.
I don't know why some people are so angry at AI/software writing with AI. It's just like being a team leader with junior(-ish) devs on the team. You don't write the code yourself, you give directions, you help/refactor/optimize where you can, that's the job.
Yes sometimes you need a team to do something, you can't code everything by yourself.
For reference, i'm not OP.
Re: Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
#155That README hits all my “this is authored by an LLM” instincts. I presume the codebase is also written by an LLM?
i thought, here on HN, we were past the "oooh it's written by a LLM it's bad!". I care about the craft, well designed systems, good clean architecture and code, etc... But i also care about reaching goals. Whether i do it working on my own, or with human coworkers or with AI coworkers doesn't matter that much to me. Yes, the result is sometimes the most important thing.
It looks like some people have a hard time accepting that software written with ai isn't a fad, it's here, it won't go away and it can be interesting and useful for the creator and for users.
Many people on the other hand have moved on and are now team leaders, except their team is mostly AIs instead of junior devs. Trade off: code is usually worse, but in the end AIs are more capable with vast knowledge and speed.
That doesn't mean we have to use AI everywhere, all the time though.
Re: Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
#156Once the tech catches up to the point that we can accurately select the right model for the task, then this ends up becoming a valuabke thing to have. You would spin it up sparingly as part of an automated discovery process maybe for 30 mins a day. And the rest of the time is spent using tiny models. I could see a future like that.
That’s not going to happen. We can’t even estimate how long it will take to complete a backlog item until after we complete it, there is no way to know how complex a task is without doing it.
We need something more sophisticated, but you could just escalate the model quality if if keeps failing the test. Not elegant but it will work.
Re: Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
#157Earlier quoted context omitted.
i thought, here on HN, we were past the "oooh it's written by a LLM it's bad!". I care about the craft, well designed systems, good clean architecture and code, etc... But i also care about reaching goals. Whether i do it working on my own, or with human coworkers or with AI coworkers doesn't matter that much to me. Yes, the result is sometimes the most important thing.
quite funny to see HN crowd downvoting this. It looks like some people have a hard time accepting that software written with ai isn't a fad, it's here, it won't go away and it can be interesting and useful for the creator and for users. Many people on the other hand have moved on and are now team leaders, except their team is mostly AIs instead of junior devs. Trade off: code is usually worse, but in the end AIs are…
It looks like some people have a hard time accepting that a readme or a documentation is not software.
Re: Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
#158Earlier quoted context omitted.
The behavior from llama-server I've seen in the past is that it fills the RAM, then completely fills the swap when the GGUF won't fit in available CPU-connected + GPU RAM. I plan to do some further testing watching iostat live and other metrics for level of constant ongoing writes to the swap, to see just how detrimental it could be to SSD write life.
You should better not use swap at all, which eliminates all problems, especially on any system that has a decent amount of DRAM. I have stopped using swap a quarter of century ago, and it was for the better. I have seen swap advocates, but I do not agree with any of their arguments. I have encountered workloads for which the amount of memory in a computer was insufficient, so the OOM was invoked, but in all such case…
Re: Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
#159Idk about you guys but I'd find 0.5t/s useless. Even for long tasks. I'd rather just shell out the money to offload as much as possible to say 2x 4060ti 16gb with tensor parallelisation. Anything but that low token rate. This is the sort of thing I'd expect in 20 years for some cyberpunk esque "turtlebot" that thinks at 0.5t/s, is solar powered and performs some menial civic maintenance background task like cutting g…
Re: Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
#160Earlier quoted context omitted.
The emdash allergy is shortcut sheep mindset. Emdashes are commonly accepted and orthographically correct punctuation in many languages. Also it looks like English might not be the author’s first language—might not lead to the best training corpus.
I speak English, I don't even know how to create one on my keyboard - is it a mac thing? I literally never saw one until chatgpt started to generate them