Earlier quoted context omitted.
Wikipedia absolutely self-corrects, that's the whole point!
it does not. it's authors corrects it. unless you see Wikipedia as the organisation and not the encyklopedia? in that case: sigh, then everything self corrects
Storm: LLM system that researches a topic and generates full-length wiki article
61–70 of 104 posts
Re: Storm: LLM system that researches a topic and generates full-length wiki article
#62At what point will it be just LLM Bots arguing with Other LLM Bots on Wikepedia edits ?
Otherwise it would be like reading HN or Reddit debates where 2 egomaniacs who are both wrong continually straw man each other with statements peppered with lies and parroted disinfo, aint got time for that.
Re: Storm: LLM system that researches a topic and generates full-length wiki article
#63I can see this being useful iif the content is generated on demand and then discarded. Publishing AI generated material is generally speaking a horrible idea and does nobody any good (at least until accuracy levels get much much better.) Even if they do it well and truthfully (which they don't) current LLMs can only summarize, digest, and restate. There is no non-transient value add. LLMs may have a place to help que…
The hardest part of good documentation is getting started. Once there are docs in place it's usually much easier to revise and correct than it would have been to write correctly by hand the first time. Think of it like automating a rough draft.
Re: Storm: LLM system that researches a topic and generates full-length wiki article
#64Earlier quoted context omitted.
That's gonna be a lot of fun to play with in a year or so. There's a concurrent explosion of 'veracity' analysis - it'll be fun to run those against wikipedia a year from now and your data. Incidentally, are you interested in mirroring your dataset and making it more robust? I'm sure I've got a few TB of storage lying around somewhere...
You can just download it yourself. Wikimedia publishes regular dumps in easily accessible formats: https://dumps.wikimedia.org/enwiki/20240320/ (the most recent for english Wikipedia)
Re: Storm: LLM system that researches a topic and generates full-length wiki article
#65Oh dear lord .... sub heading states - Storm - Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models Good luck with this storm, wiki's the world over. Just a thought but ... maybe someone should ask an org like the Internet Archive to snap-shot Wikipedia asap and label it Pre-Storm and After-Storm
Re: Storm: LLM system that researches a topic and generates full-length wiki article
#66I can see this being useful iif the content is generated on demand and then discarded. Publishing AI generated material is generally speaking a horrible idea and does nobody any good (at least until accuracy levels get much much better.) Even if they do it well and truthfully (which they don't) current LLMs can only summarize, digest, and restate. There is no non-transient value add. LLMs may have a place to help que…
current LLMs can only summarize, digest, and restate. There is no non-transient value add. Though, at a stretch, Wikipedia itself could be considered based around summarization, digesting, and restating/citing things said elsewhere, given its policy of verifiability: "Even if you are sure something is true, it must have been previously published in a reliable source before you can add it." Now, LLMs aren't well known…
Though note that there still isn't any need to publish static content. The power of LLMs is that they can be dynamic and responsive!
Even if we hypothesize that it were possible for a LLM to write a high-quality wikipedia-like output, generating the whole thing statically in advance like existing Wikipedia would be relatively pointless. It'd be much more interesting to generate arbitrary (and infinite!) pages on demand.
Re: Storm: LLM system that researches a topic and generates full-length wiki article
#67I can see this being useful iif the content is generated on demand and then discarded. Publishing AI generated material is generally speaking a horrible idea and does nobody any good (at least until accuracy levels get much much better.) Even if they do it well and truthfully (which they don't) current LLMs can only summarize, digest, and restate. There is no non-transient value add. LLMs may have a place to help que…
I think bootstrapping documentation with LLM output is a great practice. It's a wiki, people can update it from a baseline, just as long as they can see what was LLM generated to know that it shouldn't be taken as absolute truth. The hardest part of good documentation is getting started. Once there are docs in place it's usually much easier to revise and correct than it would have been to write correctly by hand the…
Re: Storm: LLM system that researches a topic and generates full-length wiki article
#68Hmm something about this title containing the word 'research' disturbs me. I associate that word with rigorous scientific methods that leads to fact based knowledge or maybe some new hypothesis, not some LLM hallucinating sources, references, quotes and all the other garbage they spit out when challenged over a point of fact. Horrifying to think peeps might turn towards these tools for factual information.
This anthropomorphism really bothers me. These tools are useful for what they’re good for, but I really dislike the agency people keep trying to give to them.
So many people think LLM means chatbot, even here on HN. So many people think agent means mentally humanoid.
But we have others, like Stable Diffusion's Web UI and Leonardo.AI - these are just tools with interfaces and the text entry for prompting is not presented as though it's a conversation between 2 people.
Someone shared an AI songmaker here recently... And there's a number of promising RAG tools for improving workflows for: Doctors, mechanics, researchers, lawyers.
I agree with you and expect the "AI character" use case to narrow significantly.
Re: Storm: LLM system that researches a topic and generates full-length wiki article
#69Oh dear lord .... sub heading states - Storm - Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models Good luck with this storm, wiki's the world over. Just a thought but ... maybe someone should ask an org like the Internet Archive to snap-shot Wikipedia asap and label it Pre-Storm and After-Storm
LLM mediocrity is just a reflection of human mediocrity, and my bet is on the average LLM to get way better much faster than the average human doing the same.
Only fine-tuned models are producing impressive work, because when we say something is impressive it by definition means not like the status quo - the model must be tuned toward some bias or other, whether it's aesthetic or otherwise, in order to stand out from the rest. And generic models like GPT or Stable Diffusion will always be generic, they won't have a bias toward certain truths - they'll be mostly unbiased which we want for general research or internet search.
So it's interesting, in order to get incredible quality of work out of AI, you have to make it specific, but in order to that, you have to train it on the work of humans. I think for this reason AI will always be ultimately behind humans, though it of course will displace a lot of work we do, which is significant.
Re: Storm: LLM system that researches a topic and generates full-length wiki article
#70I can see this being useful iif the content is generated on demand and then discarded. Publishing AI generated material is generally speaking a horrible idea and does nobody any good (at least until accuracy levels get much much better.) Even if they do it well and truthfully (which they don't) current LLMs can only summarize, digest, and restate. There is no non-transient value add. LLMs may have a place to help que…
No, you're wrong. LLMs create new experiences after deployment, either by assisting humans, or by solving tasks they can validate, such as code or game play. In fact any deployed LLM gets to be embedded in a larger system - a chat room, a code running environment, a game, a simulation, a robot or inside a company - it can learn from iterative tasks because each following iteration carries some kind of real world feedback.
Besides that, LLMs trivially learn new concepts and even new skills with a short explanation or demonstration, they can be pulled out of their training distribution and collect experiences doing new things. If OpenAI has 100M users and they consume 10K tokens/user/month, that makes for 1 trillion tokens of human-AI interaction rich with new experiences and feedback.
In the text modality LLMs have consumed most of the high quality human text, that is why all SOTA models are roughly on par, they trained on the same data. That means easy time is over, AI has caught up with all human language data. But from now on AI models need to create experiences of their own, because learning from your own mistakes is much faster. The more they get used, the more feedback and new information they collect. The environment is the teacher, not everything is written in books.
And all that text - the trillions of tokens they are going to speak to us - in turn contributes to scientific discoveries and progress, and percolate back into the next training set. LLMs have massive impact at language level on people so by extension on the physical world and culture. They have already influenced language and the arts.
LLMs can create new experiences, learn new skills, and have a significant impact through widespread deployment and interaction. There is "value add" if you look at the grand picture.