Here's a pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
Inkling: Our Open-Weights Model
221–230 of 324 posts
Re: Inkling: Our Open-Weights Model
#222> Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning. Open base models that can be fine tuned on Tinker is a great business model IMO. You (i.e. an enterprise) can own your own model & have it perform frontier-or-bette…
> that can be fine tuned on Tinker Good source to understand why this is valuable?
1.) stuff it into context
2.) figure out a way to determine what to put into context based off of what is being asked of the model (RAG)
3.) change the weights of the model to have knowledge of your data baked into it (fine tuning)
Re: Inkling: Our Open-Weights Model
#223Earlier quoted context omitted.
AllenAI is great, but they don't have the budget or remit to build large models.
I wonder if the recent sale of the Seahawks will change that. IIRC, ~$10B and all is supposed to go to charity. Not sure how much of that will go to AllenAI, though. (If any.)
Hopefully it somehow works out though!
Re: Inkling: Our Open-Weights Model
#224Earlier quoted context omitted.
I’m trying to be charitable but your comment reads as “China bad” propaganda to me. Who cares that DeepSeek and Z.ai are Chinese companies?
LLMs aren't just for coding and math. Many people understand the world through LLMs, even when it comes to philosophy and politics. If you understand the world through a Chinese LLM, you are seeing it through a biased lens stemming from biased training data. (Also, in that way, having all major LLMs developed by the US carries a risk too. We need more diversity than just the viewpoints of the US or China.)
Re: Inkling: Our Open-Weights Model
#225Earlier quoted context omitted.
I’m trying to be charitable but your comment reads as “China bad” propaganda to me. Who cares that DeepSeek and Z.ai are Chinese companies?
I think practically every government will want to put restrictions on private companies building models. Frankly the EU and the US will practically be less involved and have more pushback from the public in this than China. I think that’s less “China bad” than recognizing that China is a more authoritarian state and has far more proclivity to interfere than western states. Maybe I’m wrong? What does deep seek say abo…
Every model will have their own bias. Freedom is relative and American freedom is not the only form of freedom. China bad or China authoritarian and thus not free is a result of the red scare
Re: Inkling: Our Open-Weights Model
#226Earlier quoted context omitted.
The same business model that Deepseek is using. Open-source models + services. This is more attractive because it doesn't lock in the vendors. If I grow larger, I can decide to deploy the open-source models.
> The same business model that Deepseek is using. there is a chance their business model is absorbing government funding..
Re: Inkling: Our Open-Weights Model
#227Re: Inkling: Our Open-Weights Model
#228Earlier quoted context omitted.
Thinky has a potential answer in Tinker — give away the weights and charge for the SFT (and maybe RL down the line) to make the model more capable for specific tasks.
SFT/RL can be done without parent company.
Re: Inkling: Our Open-Weights Model
#229Earlier quoted context omitted.
> that can be fine tuned on Tinker Good source to understand why this is valuable?
If you want an LLM to have knowledge about something, the knowledge has to either exist in its weights or be provided to it in its context. Because context is expensive and limited, and models tend to get dumber the more their context is filled, there is usually more that you'd need to put into context than can reasonably fit in it in order for the model to answer questions about your data. So your options are basica…
Re: Inkling: Our Open-Weights Model
#230Earlier quoted context omitted.
NVIDIA is building Nemotron
I don't want to say this but Nemotron is not worth running on any sillicon, given Nvidia has been doing it for 3+ years, if Nvidia instead gave away GLM or KIMI API for free no one would use Nemotron the reason it's so wildly used is because Nvidia offers a Free API...