Live data from Hacker News

The deck we used to raise our seed funding

airbyte.io

51–60 of 127 posts

Re: The deck we used to raise our seed funding

#51

One thing that’s not clear to me is why is there so much competition and crowding in “data massage” space. There is Snowflake, there are all kinds of ETL tools. The customer lists these startup posts have overlaps. Is it just Marketing departments inside these companies playing around with these tools or the CIOs cycling through the hottest startup on TechCrunch list ?

They all promise to reduce your data engineering budget. The problem is that building a data connector is a one-time platform problem per data source. Once it’s solved; it’s solved. None of them solve the problem of ETL design and data warehousing design.

Re: The deck we used to raise our seed funding

#52
Having worked with Fivetran, Segment and Singer in the past I am really excited for an opensource solution like what you guys have developed. The long tail of connectors has been a real hassle when you work with mostly small companies who use very specific SaaS products.

Wish you guys best of luck

Re: The deck we used to raise our seed funding

#54

One thing that’s not clear to me is why is there so much competition and crowding in “data massage” space. There is Snowflake, there are all kinds of ETL tools. The customer lists these startup posts have overlaps. Is it just Marketing departments inside these companies playing around with these tools or the CIOs cycling through the hottest startup on TechCrunch list ?

Coming from an ad agency background I’ve seen a lot of attempts at “unifying” various data sources from client’s analytics and sales data, agency tools, and third party data sets that are all in different formats, date ranges, and scopes.

Warehousing that data might also require firewalling clients or teams for privacy or “competitive/conflict” reasons.

These aren’t difficult problems to solve with a few knowledgeable devs but that is nothing but added cost and some agencies just aren’t good at hiring the right devs - especially if their previous exposure has been basic front end web developers from their clients.

“Data warehouse” has also become a selling term even if “really big database” is a more accurate term.

Hopefully more of these companies start to distinguish themselves in this space but their competition isn’t each other - it’s entry-level data people blasting through Excel.

Re: The deck we used to raise our seed funding

#55
post #51

One thing that’s not clear to me is why is there so much competition and crowding in “data massage” space. There is Snowflake, there are all kinds of ETL tools. The customer lists these startup posts have overlaps. Is it just Marketing departments inside these companies playing around with these tools or the CIOs cycling through the hottest startup on TechCrunch list ?

They all promise to reduce your data engineering budget. The problem is that building a data connector is a one-time platform problem per data source. Once it’s solved; it’s solved. None of them solve the problem of ETL design and data warehousing design.

It sounds like you don't think solving for data connectors + necessary maintenance has a lot of value. I would agree, not FTE levels of value, but most companies I've seen in the SMB space would do well to pay $1-3k per month to have their data all housed in one spot. That lets their 1-2 DS/DE/SWE spend their time actually analyzing the data.

Maintaining connectors is also a good way to demotivate high achievers - better to have them further down the value funnel.

Re: The deck we used to raise our seed funding

#56

One thing that’s not clear to me is why is there so much competition and crowding in “data massage” space. There is Snowflake, there are all kinds of ETL tools. The customer lists these startup posts have overlaps. Is it just Marketing departments inside these companies playing around with these tools or the CIOs cycling through the hottest startup on TechCrunch list ?

It is pretty crazy.

I worked for a large organisation where management was far closer to 'technology leaders' and 'technology strategists' than engineering and data science principles and leads. They would endlessly swoop in to our division asking us to assess another product they have bought to fix the legacy problems of multiple data sources.

All of them were brittle af. They all anticipated a very idealistic data source and the absence of non-technical people curating data in excel ten different ways.

Even though we were the data science team, we usually ended up providing far more value to the organisation because we could do data engineering and cleaning and ended up being the source of truth for a lot of data required by the wider organisation. We got pitched dozens of sexy solutions to fix all our ETL problems, but when we started asking questions it was always seemed like a well designed custom pipeline couldn't be beaten for both data quality assurance, reliability and speed.

Re: The deck we used to raise our seed funding

#57
post #13

Interesting to see the competitive analysis with Fivetran in the article but then see almost identical copies of infographics used between their site and Fivetran's. Airbyte: https://airbyte.io/wp-content/uploads/2021/03/Airbyte-Seed-D... Fivetran: https://images.cms.fivetran.com/mgtdf72hs0mx/6qYtmEEotXqScar...

They launched as a Fivetran alternative. So that may not be coincidence.

Re: The deck we used to raise our seed funding

#59

Step one: join YC. Step two raise seed. Seriously, this deck would likely not have flown without the YC backing and implicit stamp of approval, once you are in YC you'd have to do pretty bad not to raise seed funding.

I wonder if there is any YC company that failed to even raise Seed Round.

Re: The deck we used to raise our seed funding

#60
post #56

One thing that’s not clear to me is why is there so much competition and crowding in “data massage” space. There is Snowflake, there are all kinds of ETL tools. The customer lists these startup posts have overlaps. Is it just Marketing departments inside these companies playing around with these tools or the CIOs cycling through the hottest startup on TechCrunch list ?

It is pretty crazy. I worked for a large organisation where management was far closer to 'technology leaders' and 'technology strategists' than engineering and data science principles and leads. They would endlessly swoop in to our division asking us to assess another product they have bought to fix the legacy problems of multiple data sources. All of them were brittle af. They all anticipated a very idealistic data…

That's exactly why we are approaching the problem with open source. It changes the dynamic of how it gets adopted. we've been in your shoes where a tool is being pushed Top-Down and now you have to deal with a super complex, super expensive, rigid & half working product.

Instead Airbyte gets adopted by engineers, data scientist... to solve one problem and then the usage expands from there. We can improve the product based on the feedback we get from the real users.

And if a feature, a connector is not there, anyone can actually add it!

Post reply on HN