Live data from Hacker News

Launch HN: Prolific (YC S19) – Quickly find high-quality survey participants

news.ycombinator.com

1–10 of 24 posts

Launch HN: Prolific (YC S19) – Quickly find high-quality survey participants

#1
Hey HN,

We’re Katia and Phelim, cofounders of Prolific (https://www.prolific.co). We help psychological and behavioral researchers quickly find participants they can trust.

We built Prolific because Katia had a hard time finding participants for her psychology studies during her PhD. She briefly used Amazon's Mechanical Turk (MTurk), but didn’t like the user experience and couldn’t get the data she wanted (UK participants). The fundamental problem we’re hoping to help with is better access to psychological and behavioral data. This is challenging in many ways: You have to balance the growth of a multi-sided platform, achieve high data quality, align incentives for all stakeholders (researchers, participants, ourselves, society), diversify the participant pool, to name some. We’re first-time founders and we’ve been bootstrapping our startup for the past 5 years during our PhDs.

Researchers build their surveys using Google Forms, Qualtrics, SurveyMonkey, Typeform, or another tool; all you need is a survey URL to get started. We verify and monitor participants so you can get data fast (most surveys are completed in We have 70,000+ survey takers in Europe and North America (for distributions of demographic variables see https://www.prolific.co/demographics) and 100s of demographic filters (try our audience checker via https://app.prolific.co/audience-checker). This means we can find many target demographics for you. For example, you can filter for Democrats vs. Republicans, old vs. young people, students vs. professionals, different ethnicities, people with health problems, Brexit voters, and even collect nationally representative samples!

Anyone can sign up as a participant and start earning a little extra cash.

It's possible to do research using existing platforms like MTurk. Actually, over 50% of behavioral research is now run online, mostly on MTurk. But there are problems with the quality of data you get from existing platforms, and worse, problems with how the people who participate get treated [1]. Our approach addresses these issues. We think the key differences are: It’s data you can trust: We mandate a minimum hourly reward of $6.50, and often rewards are even higher than that. As a result, participants feel respected and treated like valuable contributors, providing high quality data. We comply with data protection regulation and have a range of technical and behavioral checks in place to ensure high quality data [2]. Demographic prescreening is flexible and free: You can easily invite participants for follow-up studies at no extra cost. You can get niche or even nationally representative samples on-demand. Prolific is built by researchers for researchers. We try to distribute studies as evenly as possible across our participant pool through rate limiting, so have less of a problem with “professional survey takers” than MTurk.

Our bigger vision is to build tech infrastructure that empowers behavioral research on the internet. The market opportunity is significant because any individuals, businesses, and governments would benefit from better access to rigorous behavioral data when making decisions. For example, what could we do to best curb climate change? What’s the best way to change unhealthy habits? How can we reduce hate crime and political polarization? The stakes are high, and behavioral research can help us find better answers to these kinds of questions.

Moreover, although we built Prolific primarily to help academics, we've noticed that businesses have been using the platform for things like market research and idea validation. This is a new market for us that we're excited to explore. We’d love to hear about any ideas, experiences, and feedback you might have. Thank you!

[1] https://news.ycombinator.com/item?id=19719197

[2] https://blog.prolific.co/bots-and-data-quality-on-crowdsourc...

Re: Launch HN: Prolific (YC S19) – Quickly find high-quality survey participants

#2
This is definitely a space that needs innovation! How do you plan to handle the case of survey takers being real people but just skimming through the survey?

In my experience running research studies, this was the main problem with MTurk. Things like bots and blatantly junk answers were relatively rare and easy enough to detect and filter. What was much harder to deal with was the fairly high volume of users who just want to finish the survey as fast as possible. It's not so bad for short surveys, but anything over 5 minutes starts to have issues with low quality responses.

We had to introduce several control questions to check for consistency in answers and measure time to find outliers. But, it was not an ideal setup and there were still many survey takers we suspected were not paying much attention. The breakdown for our studies was something like 60% good quality answers, 35% low quality answers that are hard to distinguish from high quality, and 5% total junk answers. We ended up doing an in-person study where we got much cleaner results, presumably because people pay better attention when they feel someone is watching them.

Wages don't seem to do much for this. We were paying a relatively generous $4 per HIT, estimating 30 minutes per HIT when in reality in reality the average time was under 20.

So I'm wondering, are you able to share with us how exactly the "unusual data patterns" and "technical and behavioral checks" can help ensure quality in an ecosystem where users are generally motivated to 1) finish surveys as fast as possible, and 2) appear as if they are giving high quality responses when they are not?

Re: Launch HN: Prolific (YC S19) – Quickly find high-quality survey participants

#3
Very cool!

I signed up and will be looking out for surveys. Looking at the demographics, I noticed the survey panel is similar to that of Hacker News: white English speaking 20-30 year olds many who are in school. I'd love to hear about your efforts to have a more nationally/globally representative sample?

Re: Launch HN: Prolific (YC S19) – Quickly find high-quality survey participants

#4
I'm locked in to my current vendor, but I just wanted to share some of my thoughts about some pain points I saw in your sales pitch.

My experience is that most academics pay less than $6.50 per hour for an initial point of contact. For instance, I am currently fielding a survey (N > 5,000 per week, cross-sectional rather than panel) and we pay about $2 all-in, including the provider's charge, for a survey that's about 20 minutes. We'd fall afoul of your compensation rate pretty substantially. If we wanted to do some panel work and we needed re-contact, we'd definitely ramp up our payment quickly to help avoid attrition, but for the first contact, no.

If we wanted to be paying out your rate, we would almost certainly have to get an additional sponsor partner to piggyback some consumer research on top of our actual treatment. We don't want to do this, it's hard enough dealing with our primary funders. This makes me believe you are mostly targeting commercial / market behavior researchers. That's fine, but the pitch suggests you want academics. For transparency's sake, what is your balance of private and university clients as of today?

Second, even working with large sample providers, their pools are often fairly small. We requested 5,000 unique respondents a week for a year and found our sample provider could only guarantee a 1-2 month lockout. Obviously the effective pool you need to guarantee 5,000 * 52 is enormous and so we were expecting to have to negotiate on lockout, but all of this dances around the fact that sample providers are not transparent about the size of their pool and researchers like us are constantly worried about fraud both by sample providers and by respondents. How large is your pool?

Finally, this kind of quota sampling relies on our ability to weight the sample to the population. Weighting is totally permissible, but responsible weighting is going to cap the weight at the high end -- no one wants the one black Republican to skew the entire poll because they have a 150 weight on their observation (this isn't me getting needlessly political, this happened with the USC tracking poll last election cycle). In my experience, the hardest thing about quota sampling as opposed to the old RDD phone samples 20 years ago is that it's very difficult to get high education / high SES / high income respondents. High income respondents should be 10% of the population and they simply are nowhere near 10% of the pools that you get from standard recruitment methods. Can you speak a bit about a) how you recruit high income people into your pool; and b) what percentage of your pool would be high income (say HHI > $125k a year or so).

Finally, what is your pool attrition rate? If someone takes their first Prolific survey today, what is the probability they will still be taking a survey a year from now? It's nice that you have re-contact as part of your system, but the problem in my experience is not figuring out how to recontact, it's getting people to stay engaged for a long time.

Hope you have good answers to these questions, and that if you do, answering them here will help you get positive exposure from other readers.

Re: Launch HN: Prolific (YC S19) – Quickly find high-quality survey participants

#5
post #2

This is definitely a space that needs innovation! How do you plan to handle the case of survey takers being real people but just skimming through the survey? In my experience running research studies, this was the main problem with MTurk. Things like bots and blatantly junk answers were relatively rare and easy enough to detect and filter. What was much harder to deal with was the fairly high volume of users who just…

It’s a great question and a tough problem. We have written a little bit about the problem of ‘slackers’ (participants with low attentiveness and engagement) [1,2]. The things we’re doing right now to address this problem is to 1) test for attentiveness and engagement before participants take part in real studies, 2) distribute surveys more broadly to reduce the % of “professional survey takers” in each study and 3) educate researchers about ways to use good attention and engagement tests [3] in their studies. We can then feed this data back into our system so we can iteratively improve on data quality.

I can’t go into too much detail about the attention and engagement checks we have built in, but if you sign up as a participant and pay attention you might spot some. In the long run, we think good feedback systems and high trust will be key so that our pool iteratively gets better and participants don’t feel as incentivized to cheat. It’s key to make sure that participants feel that their high effort responses get fairly rewarded (both financially and with social appreciation).

[1] https://blog.prolific.co/how-to-improve-your-data-quality/

[2] https://blog.prolific.co/bots-and-data-quality-on-crowdsourc...

[3] https://researcher-help.prolific.co/hc/en-gb/categories/3600...

Re: Launch HN: Prolific (YC S19) – Quickly find high-quality survey participants

#6
post #3

Very cool! I signed up and will be looking out for surveys. Looking at the demographics, I noticed the survey panel is similar to that of Hacker News: white English speaking 20-30 year olds many who are in school. I'd love to hear about your efforts to have a more nationally/globally representative sample?

Hi, Katia here. Thanks for your question!

Right now we’re available to participants in OECD countries and it’s probably going to be another while before we can expand globally. Within the 36 countries that we’re in, we’re currently trying to work out what incentives might work best: How can we encourage the average citizen to sign up and take part in online research? What about very hard-to-reach demographics like professionals?

We feel that cash generally works well as an incentive, but it’s not everything. A lot of the time people want to contribute to projects they personally care about. So on Prolific’s end, it will come down to matching projects with the right participants. Another type of incentive could be to offer to participants that Prolific donates their earnings to charities on their behalf, and perhaps Prolific could match their donation?

We’ve recently launched quota-based representative samples for the US and UK [1], where we stratify based on sex, ethnicity, and age. A by-product of this feature is that it specifically invites niche demographics to participate, helping fill gaps. We hope that having more surveys available for demographics that are in the minority (e.g. certain ethnicities and age groups) will improve our ability to recruit these participants (although there’s a bit of a chicken-and-egg problem here).

Another thing we’d like to do is launch a mobile app for participants, so anyone can take a quick survey on the go (while waiting in line, chilling at home, commuting). Making our surveys more accessible to participants through different channels should help represent more people in society.

And then there’s user experience. We’re working on making our site as self-explanatory as possible, so anybody can sign up and start participating, even if you’re someone who doesn’t spend much time on the internet.

What do you think––what might be other ways to diversify our participant pool? Very curious

[1] https://researcher-help.prolific.co/hc/en-gb/articles/360019...

Re: Launch HN: Prolific (YC S19) – Quickly find high-quality survey participants

#7

I'm locked in to my current vendor, but I just wanted to share some of my thoughts about some pain points I saw in your sales pitch. My experience is that most academics pay less than $6.50 per hour for an initial point of contact. For instance, I am currently fielding a survey (N > 5,000 per week, cross-sectional rather than panel) and we pay about $2 all-in, including the provider's charge, for a survey that's abou…

> what is your balance of private and university clients as of today?

Our clients are about 80-90% academic, and we find that there’s a strong move within academia towards fair payments. In the long term we think that this will show in the data quality and reliability of samples such that paying fairly will result in the best value. Regarding cost, we're about $2 total for a 15min survey right now.

> How large is your pool?

Lack of transparency about pool size is one of the most frustrating things about online panels. This is one of the reasons why we don’t report our total pool size, but only our ‘active and accessible’ pool (participants who have been active within the past 90days). As a result, you can expect > 50% of the numbers of eligible participants we report to actually take part in the next 24/48hours. We have ~70,000 active participants (20,000+ in the US). I expect we would be able to get a sample of 5,000 people in > a) how you recruit high income people into your pool;

To be honest, we haven’t cracked this nut yet and we expect our pool to be unrepresentative for high income people also right now. Ideas we have to help address this problem are to 1) introduce charitable donations for those who aren’t motivated by cash incentives 2) improve non-financial incentives (e.g. feedback on the impact your data is having) and 3) have highly targeted invites, we hypothesize high earners would be willing to help out if we need someone of their demographics in particular, and if this was communicated well. If you have suggestions, we’re all ears!

> and b) what percentage of your pool would be high income (say HHI > $125k a year or so).

According to self report we have ~2,000 participants from a HHI (>$100k/year on our screener), though it’s possible this suffers from slight inflation and we don’t (yet) have a way to verify income levels.

> Finally, what is your pool attrition rate?

Our annual retention rate is ~40% or so (e.g. if a participant takes part in a study, they’re about ~40% likely to do so again 12months later). There’s a balance between ‘refreshing’ our pool and keeping high engagement, and we’re working on keeping “naivety” high while allowing for studies over a long period!

Re: Launch HN: Prolific (YC S19) – Quickly find high-quality survey participants

#10

We use Survey Monkey. This seems similar/identical. How do you compare to that? Is there a reason to pick you over them? They charge about $1/response.

The short answer is that with Prolific you can use any survey/experimental software you'd like. Behavioural researchers (our primary customers) don't tend to use SurveyMonkey, but instead use Qualtrics, Gorilla, and custom experimental software which have many features important for behavioural research (e.g. randomisation and reaction time recording).

In addition, SurveyMonkey audience uses a survey exchange provider CINT [1] and as a result you don't have any transparency of where exactly the participants are coming from, how they're vetted, the user experience they're having, and you can't communicate directly with participants or retarget them for longitudinal studies. All of which is important for research and are features that Prolific provide.

[1] https://www.cint.com/press-release/surveymonkey-audience-exp...

Post reply on HN