Earlier quoted context omitted.
I’ve had a similar experience where I could email PMs and we even had Google forward deployed engineers on-site. The trouble is that their customer service is terrible. We reported numerous issues and there was no tracking at all. For one UI bug in their webapp, a PM completely disregarded the reproduction instructions I gave him and made me have a half-hour kong Hangout with a remote engineer to prove the bug existe…
Cloud Dataflow is also vastly more expensive than using pyspark and Apache Beam is 10 times slower for most operations than Flink or Spark. I have to try the new flexible scheduling but DataFlow the last time I tried it 2 years ago was a massive rip off. Also DataPrep was amazing until you run the DataFlow pipeline it produces and a Have had good experiences once the data is in bigquery and bq keeps costs low. But la…
Do you need to run this pipeline more than once a day? Is it sufficiently important for your business case? I feel like 25 is a very small sum if any of them are remotely true.
On a related note, how practically reliable is the DataPrep->DataFlow workflow? Could a reasonably smart analyst with zero programming experience set them up and run them? Feel like that's who this workflow is built for, but I've not had the best of experiences with Trifacta demos in AWS with reliability.