Live data from Hacker News

Show HN: DataFuel.dev – Turn websites into LLM-ready data

datafuel.dev

11–20 of 37 posts

Re: Show HN: DataFuel.dev – Turn websites into LLM-ready data

#11
I thought this might be interesting to share and potentially useful for the author of Datafuel as a comparison. I recently built something similar for a small app [1].

I use Bun.js's fetch to crawl pages, process them with Mozilla’s Readability (via JSDOM), and convert the cleaned content to Markdown using Turndown. I also strip href attributes from links since they’re unnecessary for my use case, and I don't recurse links. My implementation is basic, with minimal error handling and pretty dumb content trimming to stay within prompt tokens limit which could use improvements! I also found this Python library that seems a lot fancier than what I need, but also a lot more powerful [2].

I’m curious where a solution like Datafuel excels, especially since it already has customers? From the top of my head, the real complexity in scraping appears when processing a sizable number or URLS regularly and becomes more of a background processing / scheduling problem.

I feel like something like Datafuel could become more adopted if it was a nicely put together as a library to crawl locally, and then if you find yourself crawling regularly and want to delegate the scheduling of those crawls, you could buy into the service: "ping me back when these 10_000 URLs are done crawling", or something like that.

--

1: https://github.com/EmmanuelOga/plangs2/blob/main/packages/ai...

2: https://github.com/adbar/trafilatura

Re: Show HN: DataFuel.dev – Turn websites into LLM-ready data

#12

This is a pretty crowded market, e.g. Brightdata (most feature complete), Firecrawl (focused on api and sdk), etc

yes exactly working on building a competitive advantage, but the AI space is so big and only getting bigger.

Any feedback that could help datafuel becomes more unique?

Re: Show HN: DataFuel.dev – Turn websites into LLM-ready data

#13
I am interested, but why should I use this one over jina ai reader (which is also free) or firecrawl, or the ten other puppeteer + readability + turndown pipeline (or even a AWS lambda doing the same) ? This is not sarcastic I am genuinely looking for something fresh in the field.

Re: Show HN: DataFuel.dev – Turn websites into LLM-ready data

#14
post #4
post #2

Great, congrats on your launch. 1. Does it take care of Bot detection. Most sites will have it. 2. Is this something similar to Firecrawl - https://www.firecrawl.dev/

Yes, it has an extensive proxy IP and retry system in place to bypass bot detection. I’m also trying to gather more feedback to identify the killer feature: - Adding vectorization to Pinecone out of the box? - Adding multiple integrations like n8n, etc.? Any crucial pain points to avoid?

Are you concerned about making a product that does this? The legal aspect of accessing a computer system that is intending to block your use seems worrisome.

Re: Show HN: DataFuel.dev – Turn websites into LLM-ready data

#15
post #9
post #6

Earlier quoted context omitted.

It boggles my mind that you would launch without that as a prime directive.

OP just graciously accepted that feedback, no need to be condescending :)

It's kind of tone deaf to launch a tool like this without considering this in the current climate. Not a popular take on hackernews but everyone outside the tech space is pretty pissed about this stuff.

Re: Show HN: DataFuel.dev – Turn websites into LLM-ready data

#16
post #13

I am interested, but why should I use this one over jina ai reader (which is also free) or firecrawl, or the ten other puppeteer + readability + turndown pipeline (or even a AWS lambda doing the same) ? This is not sarcastic I am genuinely looking for something fresh in the field.

do you need to embed it directly in pinecone ?

If yes then DataFuel is the right choice. Adding this feature as we speak.

Please let me know :)

Re: Show HN: DataFuel.dev – Turn websites into LLM-ready data

#17
post #8

Will this benefit sites or internal wikis which have well written content, good search and SEO? I interviewed at a few companies which apparently enables managers to use AI as an excuse to implement text search.

I guess so if you goal is to have people knows about your content, might have a small SEO bump

Re: Show HN: DataFuel.dev – Turn websites into LLM-ready data

#18

I thought this might be interesting to share and potentially useful for the author of Datafuel as a comparison. I recently built something similar for a small app [1]. I use Bun.js's fetch to crawl pages, process them with Mozilla’s Readability (via JSDOM), and convert the cleaned content to Markdown using Turndown. I also strip href attributes from links since they’re unnecessary for my use case, and I don't recurse…

yes exactly,

The main issue in scraping:

- If you scrape a lot, you will be block based on you IP; You need to use PROXY - Scraping entire website need specific logic, retries and more - It becomes an heavy background job

All the above takes time, so if in your business it is not your core feature, likely better to outsource it.

Good job doing it tho!

Re: Show HN: DataFuel.dev – Turn websites into LLM-ready data

#19
post #4

Earlier quoted context omitted.

Yes, it has an extensive proxy IP and retry system in place to bypass bot detection. I’m also trying to gather more feedback to identify the killer feature: - Adding vectorization to Pinecone out of the box? - Adding multiple integrations like n8n, etc.? Any crucial pain points to avoid?

Are you concerned about making a product that does this? The legal aspect of accessing a computer system that is intending to block your use seems worrisome.

It is the responsibility of the user. Everyone should be responsible for their own actions. We still allow knives to be sold, and most people use them for good.

Re: Show HN: DataFuel.dev – Turn websites into LLM-ready data

#20
post #16
post #13

I am interested, but why should I use this one over jina ai reader (which is also free) or firecrawl, or the ten other puppeteer + readability + turndown pipeline (or even a AWS lambda doing the same) ? This is not sarcastic I am genuinely looking for something fresh in the field.

do you need to embed it directly in pinecone ? If yes then DataFuel is the right choice. Adding this feature as we speak. Please let me know :)

Interesting but we process documents before embedding them, and have specific requirements for the embedder.

Having developed a couple of page to markdown myself, I think the bigger challenge is to make sense of so many pages that rely on spacial organisation of information that only makes sense to human, or even presence of images. One way to do it is to render the page as an image and extract data with a vision llm. But you do need heuristic on when to do classic extraction and when to use vision, plus get rid of cookie banner and overlays. This is more complex and costly, but have real business value, for the one that can pull it off.

Post reply on HN