I've been vary of using ai to search considering all the spam out there. I think I'd rather, perhaps naively, whitelist wikipedia, reddit, arxiv, some news sources, etc than include everything. Is there nothing out there that does this? I'm paying for kagi and I can see that it has an api, is that maybe sufficient if configured properly?
Three sites made 215,128 “best software” pages for AI. Perplexity cites them
91–100 of 265 posts
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#92Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#93Earlier quoted context omitted.
> asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored [...] It always picks its own ...is not the same as claiming... > LLMs favor LLM-generated passages over human written ones Here, you're using the same LLM to both produce and judge the resulting work. If anything, I would expect an LLM to tend to prefer its own work given that the same training is produc…
It's not intuitive to me for why preference for its own writing would emerge, and during what type of training or tuning. Perhaps something like: learning to identify what source files it has worked on by the code style alone, because tasks may give human code (public repos, etc) and ask to make changes.
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#94I used one of the 12-month free Perplexity offers when they were everywhere. It felt slightly useful at first for simple queries where I didn’t want to go through the top 10 Google results manually. If I was looking for a specific recipe I remembered or a help page or user manual it would usually find it quickly. Then they started optimizing for speed of responses over quality of results. I can enter a query and see…
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#95Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#96If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated we…
OMG if this is true, do you realize what this means? The easiest way to do AI SEO is to generate all your content with AI, and we've seen what SEO does to the web... The Internet is doomed. Time to start some human-only darknets.
> Time to start some human-only darknets.
I know very little about darknets. How could you ensure that they are human-only?
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#97It's difficult to read more than a few sentences, when this itself is clearly a Claude artifact.
Agreed, I thought the subject matter was interesting enough to try and labour through the tedious prose, but once I got to "Their scale is the point." I just had to stop and just skim the rest.
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#98None of what we're doing with tech these days is something we should be doing.
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#99Earlier quoted context omitted.
If you send people to the optimal website containing exactly the information they are after, then you get fewer ad impressions than if you send them to a suboptimal website that has them going back and clicking on more links.
This is like suggesting you can show people more ads by keeping them in line at the DMV for longer. Try that at your peril. People aren't at the DMV to waste time and there's a reason the DMV is hated.
This is just another symptom of a lack of antitrust enforcement.
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#100Earlier quoted context omitted.
And it's all because of ads. The incentives in an ad-funded internet are just always going to lead to this sort of thing. The most important thing is getting the user to load your page, not actually satisfying their query. Let's hope the LLM model continues to be paying for credits, because any that move to ad revenue will become useless for real work.
> And it's all because of ads. Not really. A while ago there was a news piece stating that Israel was behind a series of fake think-tanks with very accessible websites which were created with the express purpose of feeding AI agents with alternative facts aligned with their foreign policy. If anyone has the link at hand, please post it.
People with an axe to grind or states with an agenda are already devoting tremendous effort toward affecting LLM models and it is very difficult to determine real from astroturf for humans let alone an LLM trying to train.
Much like PageRank now that the cat's out of the bag all the current approaches may prove to be useless in the long run.