Live data from Hacker News

Be good-argument-driven, not data-driven

twitchard.github.io

81–90 of 168 posts

Re: Be good-argument-driven, not data-driven

#81

Earlier quoted context omitted.

Why? Inadequately “technical”?

I think there are lots of reasons why. One possible reason: no one whose job it is to write Python scripts was ever promoted for making an Excel spreadsheet when that is the simpler and more practical approach. And no manager of people who write Python scripts is going to be able to use that Excel spreadsheet to sell "I need more responsibility and head count." People tend to follow incentives, rather than focusing o…

Excel has a history of forced format updates, breaking incompatibility. I know people who banned it because they got tired of marching to MSs upgrade beat.

Python 2 to 3 upgrade aside, can’t really say the same about the language.

There are a number of good arguments out there that might violate an engineers perception, which one might call a cognitive data model built through training and experience.

There is no theory that makes any given engineering path “wiser” than others. Just engineers chasing incentives to be engineers.

Re: Be good-argument-driven, not data-driven

#82
post #42
post #35

The related problem that I see actually more often is the "you don't have big data" problem. You know, in data science, you see people spending hours writing pandas scripts that replicate a few clicks in excel for a one of analysis. You see datasets of a few gigabytes being processed with spark when SQL would be fine. You see ML techniques being thrown at questions that could be answered simply and reliably with basi…

Talking to people is not going to help you either. You end up getting a lot of noise and making sense of what you hear is difficult. When you keep probing you will get to hear stuff thats not really critical and just often made up because you ask too many questions. Classical trap of market research.

ycombinator startup school disagrees and says it's one of the two CRITICAL things founders must have a hand in.

Of course you need to interpret it but its incredibly important and I do not think you really know what you are talking about.

https://www.ycombinator.com/library/6g-how-to-talk-to-users

Almost all the major fails I have seen in my career have been some derivative of not understanding your users.

Re: Be good-argument-driven, not data-driven

#83
post #42
post #35

The related problem that I see actually more often is the "you don't have big data" problem. You know, in data science, you see people spending hours writing pandas scripts that replicate a few clicks in excel for a one of analysis. You see datasets of a few gigabytes being processed with spark when SQL would be fine. You see ML techniques being thrown at questions that could be answered simply and reliably with basi…

Talking to people is not going to help you either. You end up getting a lot of noise and making sense of what you hear is difficult. When you keep probing you will get to hear stuff thats not really critical and just often made up because you ask too many questions. Classical trap of market research.

You have to do both. You can't just look at data & you can't just talk to users / customers without looking back at data.

Re: Be good-argument-driven, not data-driven

#84
post #35

The related problem that I see actually more often is the "you don't have big data" problem. You know, in data science, you see people spending hours writing pandas scripts that replicate a few clicks in excel for a one of analysis. You see datasets of a few gigabytes being processed with spark when SQL would be fine. You see ML techniques being thrown at questions that could be answered simply and reliably with basi…

frankly, there’s only a tiny handful of these mythical saas “1000 users each paying 1 mullion dollars” companies. the vast, vast majority of saas startups are serving millions of “users” - i put that in quotes because these aren’t real users or customers. they are real people checking out your product - but they aren’t users or customers.

if you set up a gas station near the off ramp of some major interstate, say I-65 North, you will see cars pulling in to fill up on gas. maybe buying a coffee. now, these aren’t your customers in the traditional sense of a Target or Walmart customer. Because you will never see them again. They were driving from town A to town B via the interstate- they started running out of gas and needed to refuel, so they are in your gas station now. Once they gas up, off they go. They aren’t going to come back to you and establish a customer relationship or something. We’ve all been to tons of gas stations on the interstate and we’ll probably never go back to the same one twice - unless we are plying the same route everyday like a truck driver. So the task is to find and convert these truck drivers, who are the true repeat customers.

I was working on an android app which had like millions of unique cookies. When they hired me they said we have million of users. No you don’t. If you put out an android app in some popular domain, say news, entertainment, tax accounting etc- people will download and “use” your app. they are checking it out. they aren’t users, in the sense they aren’t using it everyday or want to have a relationship with you, pay subscription etc. conversion stats are minuscule, like 0.01%. So maybe 1 out of 10000 users is the truck driver. The vast majority will never ever use your app again. To do data science with these millions of rows of user interactions and find some nuggets just because you know your way around pandas or sklearn is a fool’s pursuit. To ask foolish questions of your data, like why are all these people churning, is silly - they aren’t your users, they haven’t converted, they are just checking it out. In that sense, its a waste of time and resources to do so much data crunching. Look at actual conversions, which are probably a few thousand people, not millions. Reach out to those thousands and maybe a few tens will give feedback and then continue to iterate on the product based on that.

Re: Be good-argument-driven, not data-driven

#85
1. If there are no good arguments in the collective - there's no retrospective and it's primarily a management and psychological issue. No one is able to fully self-reflect and it breaks the existing delegation / escalation chains, respectively.

2. If there are no viable data sources, when it can be proven that there's a correlation with an actual business processes, - it's a management problem. People Can't establish viable metrics, once again, mostly due to 1.

This is something any company of any size and any budget can struggle with due to lack of XP and the usual collective XP-accumulation / knowledge sharing deficiency. You can't self-reflect onto something you haven't learned about, yet. And due to 1 this is a closed loop because lack of XP can't be escalated accordingly, most of the time it's also a Workplace Deviance factor.

3. Practically, it ends up in a bouquet of Workplace Deviance because no one in the end will be willing to take the blame and actual responsibility to fix anything.

Any Problem vs Solution type of culture will worsen things a lot i.e. "All the blame and no Compassion". Companies are usually forced to adopt some Teal stuff in the end, maybe for really no other good reason, but just to keep on growing.

The idea of hiring HR that can "work by the booK" and actually build up a personal profile of how anyone could fit into all this mess is impossible by definition - due to Employee Silence and broken retro no one will be willing to expose all the shit that is happening, in the first place... So, most of the time I see Kitchen Sink companies with volatile outcomes where there really no one who could even be able to listen to any arguments, in the first place.

Google's internal ML-driven productivity metrics became a meme already for all the reasons described above. You can't reason with Toxic and Inadequate people.

Also Asana claim that Social Loafing is a myth and everything else is a retro deficiency really wrong - retro can prevent and display certain glorious occasions, but it's not a root cause of any psychological effect by definition.

Re: Be good-argument-driven, not data-driven

#86

Earlier quoted context omitted.

Unfortunately, your reality-driven approach has ~zero emotional appeal for most managers, exec's, and alpha-data-scientist wanna-be's.

Why? Inadequately “technical”?

- "You just talked to them and concluded this? What certainty you can have on this conclusion, and how can we trust you just didn't want it to be true from the start?"

A few slides showing the data, a boring 10 minutes about methodology, and finally the conclusion brings an air of reliability that you can't replicate for knowledge instead of data.

Re: Be good-argument-driven, not data-driven

#87
post #72
post #35

The related problem that I see actually more often is the "you don't have big data" problem. You know, in data science, you see people spending hours writing pandas scripts that replicate a few clicks in excel for a one of analysis. You see datasets of a few gigabytes being processed with spark when SQL would be fine. You see ML techniques being thrown at questions that could be answered simply and reliably with basi…

Would you say the big data threshold moves every year? That would explain why people think a <1TB is big data.

>Would you say the big data threshold moves every year?

It moves with Moore's law. Big data is anything that cannot reasonably fit into memory for a single server, so yes that number is well over 1TB now.

Re: Be good-argument-driven, not data-driven

#88

This reminds me a lot of the discussion of the scientific method by Karl Popper, and David Deutsch who was very influenced by Popper. "Being data-driven" sounds very empirical . Just look at the data, and see what you find in it. But you can't just let the data "speak for itself" without an explanation or a theory that interprets the data. Popper in Conjectures and Refutations : > Observation is always selective. It…

Doesn’t the scientific method specifically say you can’t start with the data, you have to start with a hypothesis otherwise you are subject to all sorts of selection/hindsight biases. I mean you can start with data, but then you have to develop a hypothesis and use that to create an experiment that generates new data in order to reach a conclusion. It seems like that is the compromise the author is looking for, start…

The scientific method as taught in K–12 schools is largely pablum. Often, the real process (beyond iterating off prior research) begins with collecting data, then noticing patterns to make a hypothesis to be tested with targeted data collection.

Re: Be good-argument-driven, not data-driven

#90
post #35

The related problem that I see actually more often is the "you don't have big data" problem. You know, in data science, you see people spending hours writing pandas scripts that replicate a few clicks in excel for a one of analysis. You see datasets of a few gigabytes being processed with spark when SQL would be fine. You see ML techniques being thrown at questions that could be answered simply and reliably with basi…

Unfortunately, your reality-driven approach has ~zero emotional appeal for most managers, exec's, and alpha-data-scientist wanna-be's.

Data has CYA appeal.
Post reply on HN