Live data from Hacker News

Hundreds of extreme self-citing scientists revealed in new database

nature.com

201–210 of 215 posts

Re: Hundreds of extreme self-citing scientists revealed in new database

#201
post #21

The core problem here is that universities think that citation statistics are a useful metric to evaluate the quality of the work of a scientist. There's plenty of evidence that this is not the case or that even the reverse may be the case [1], but this idea refuses to die. [1]

It sucks as a metric but it does have some rough correlation in most cases, and I'm not aware of any better easily measurable metric - if you have one in mind, it'd be great to hear. The alternative of having a bureaucrat "simply judge quality" IMHO is even worse, even less objective, and even more prone to being gamed. The main problem is that there is an objective need (or desire?) by various stakeholders to have s…

>I'm not aware of any better easily measurable metric

Why should an easily measurable metric which has meaningful value exist? It doesn't seem obvious to me that it should at all. Determining the capability of a researcher is inherently a very complex intellectual task. The desire is to reduce that task to something which removes the need for the person doing the evaluation to read and understand the produced research, or to even understand the field of study in many cases. Perhaps, instead, those who are put in charge of things like awarding grant funding, granting tenure at universities, and deciding who to hire to teach ought to be expected and required to evaluate the research on its merits. This would greatly increase the intellectual sophistication and capability needed for people in those positions, but the alternative will always be fairly easily exploitable because it is easier to goose a metric than to do solid research.

We see the shortcomings of trying to reduce complex intellectual challenges to checklists or metrics all the time. And we simply ignore the alternative of relying upon intellectually capable people meeting the challenge. Personally, I don't understand why.

Re: Hundreds of extreme self-citing scientists revealed in new database

#202
post #196
post #168

Earlier quoted context omitted.

This probably varies by field, but the "large project" thing can be gamed too. So, for example, in biomedicine you often have lots of people on a paper who might only read a draft, make some trivial suggestions, and then be added as an author. As a result, there's this pressure for large groups to form, where everyone is added and everyone can cite each other. This doesn't mean the projects are bad, but it does lead…

The problem is with the binary attribution. Either you're an author of everything in the paper or you're not an author at all. Software world solved this issue with version control systems like git. And if scientists write papers in latex or other text-based formats it's trivial to use version control system for that too. Then when you quote a fragment you do "git blame" on it, and you see who created and edited this…

Except that's not how science is done. Some people are bad writers and just don't touch the paper at all. Sometimes you have a grad student do all the work for a paper and someone else writes it, or the majority (common in the first year or two, or with undergrads). Just because they had no or little commits should they not be the main author? They did the work after all.

Another example, my advisor doesn't like git. While writing papers I and my collaborators use git but send an email copy to my advisor. Clearly he's going to be on the paper because he's my advisor but you'll see zero commits from him.

I think it's just too easy to think that technology solves this in a trivial way. It's complicated. You have people from different eras working on things. And this is in a CS program, mind you. In different fields it gets much worse very quick.

Side note: go look at papers from top tier universities. You'll notice that they frequently cite colleagues at their University. Is this because they are gaming the system? Is it because they are doing the most related research (which is VERY common for a single University to work close)? Or is it a combination. In all likelihood it's a combination because citations matter. The h index is used in your performance because this is meant to be how impactful your paper is, but the system can definitely be manipulated (and likely isn't happening for malicious reasons nor necessarily unethical reasons)

Re: Hundreds of extreme self-citing scientists revealed in new database

#203
post #170

Earlier quoted context omitted.

I don't disagree with what your saying, but I think LSST's builder concept is actually quite amazing and the opposite side of that coin. For people building the telescope (think hardware, software, logistics, everything before the science can be done), many of whom are not academics, and don't typically get authorship or write papers, it's great to get credit for working on the project in a formal, public way. You do…

The builder concept is actually really appealing. Academia can tend towards the same problem as consulting, where tasks get sharply split between "credit-producing" and "not worth doing". Answering questions about your past papers; looking over someone else's proposed methodology; or cleaning up an internal tool into one you can share are all great tasks for advancing the field, but none of them bolster a CV, earn gr…

This definitely happens on LIGO. You have hundreds of authors. My optics professor in undergrad was never a first author but he sure is an author on a lot of papers.

Re: Hundreds of extreme self-citing scientists revealed in new database

#204

Earlier quoted context omitted.

No, the answer is that there is a limited number of scientists and a limitless number of research directions. This doesn't have to be correlated with brilliance. In fact, it can be easier to research some of the less popular paths because there is less competition and more low-hanging fruits.

Fine. The more accurate word would be "advanced" instead of brilliant.

You're arguing semantics like a Humanty's major. Your exact definition of "advanced" or "brilliant" have barely any relevance to the topic at hand.

Re: Hundreds of extreme self-citing scientists revealed in new database

#205

Earlier quoted context omitted.

Just to be clear, I understood what you were saying before. I just disagree. I think the right approach is to select trusted experts and let them make the decisions about their fields. Again, since this is a tech community, let me use that for an analogy. It's a classic problem for non-technical founders to evaluate their technical hires. They aren't qualified. The right solution is not to find some gameable metric o…

It seems like "Selecting trusted experts" alone would defer more to human subjectivity and biases than would be necessary if objective measures were utilized as much as possible. Existing community/expertise based moderation and reputation systems might not be directly transferable or adequate. But it shows there are new ways to think about more decentralized measures of reputation that are new to this century and ha…

I understand why it seems that nominally objective measures would be better. But I don't think cross-field, non-gameable objective measures of research quality are practically possible

I also don't think it's a problem that different groups have different interests, etc. As I say elsewhere, I think that diversity is the solution.

Re: Hundreds of extreme self-citing scientists revealed in new database

#206
Having conducted a reasonable amount of academic and scientific research, this metric is more likely to be mischaracterizing research than revealing any issues. This doesn't even establish a causal-link between self-citation and poor research quality, it just assumes it.

Most researchers continue to do new research on the same concept after a publication, and they will of course site their earlier work when continuing. Additionally, post-graduate researchers often have their names placed on the research of grad students they are in charge of, even though they often have minimal involvement in the research or conclusions drawn.

You might be able tell something from the ratio of other authors from all citations to the number of self-citations, but only if you could eliminate self citations that were not either inclusion by proxy or cases where they are merely continuing research on the same topic with new methodologies.

There are already methods for identifying bad research, none of which can be achieved through the use of non-human-assisted data analysis of the authors list of research. The only way to be sure is critical review and 3rd party verification of results with repeated experiments.

Re: Hundreds of extreme self-citing scientists revealed in new database

#207

Earlier quoted context omitted.

The purpose of PHDs are to move human knowledge forward. You have to do an analysis of something that, in all likelihood, nobody has done before (or not enough to be considered settled).

But then your analysis has to be challenged as well. And the challenges should be published. Success or fail. If you live in your own bubble the needle doesn't move forward.

[deleted]

Re: Hundreds of extreme self-citing scientists revealed in new database

#208
post #168

Earlier quoted context omitted.

This probably varies by field, but the "large project" thing can be gamed too. So, for example, in biomedicine you often have lots of people on a paper who might only read a draft, make some trivial suggestions, and then be added as an author. As a result, there's this pressure for large groups to form, where everyone is added and everyone can cite each other. This doesn't mean the projects are bad, but it does lead…

> Science gets done but the rewards seem to filter preferentially to those who are able to game the system, and the system exists out of a need to make one's self look as productive as possible This is largely due to the current model of science funding, and not just in the US. Here in Germany, many involved in public (i.e. at an university, not inhouse r&d at a company) science only get 1-year-limited chain contract…

I agree that the root cause of a lot of it (although not all of it) is funding. I basically share your perspective about long-term funding. I personally would like to see indirect costs in the US eliminated, or much more heavily reduced, audited, and justified. I also think there needs to be dramatic shifts in funding along the lines of what you mention. Proposals have already been floated, by former federal funding heads no less, along these lines. Lotteries would be good, as would awards based on merit rather than application (Hungary's model of funding people based on bibliometric factors is a good idea, even if it runs into the problem of gaming citations as pointed out here). Longer-term funds also seem like a good idea, which is the whole idea of tenure in theory.

There are other factors at play too, that are harder for me to pin down. Funding models are a big problem, but there's something related to attention-seeking or metrification at play too. Some of this has probably always been around, but in talking to older colleagues I get the sense that things are much more splashy and fad-driven than they used to be, with much greater pressure to produce in volume. A colleague explained that when it's that much easier to write and publish a paper, there's more of an expectation that you do more of them, even though the idea development time isn't any shorter.

Re: Hundreds of extreme self-citing scientists revealed in new database

#209

Earlier quoted context omitted.

It seems like "Selecting trusted experts" alone would defer more to human subjectivity and biases than would be necessary if objective measures were utilized as much as possible. Existing community/expertise based moderation and reputation systems might not be directly transferable or adequate. But it shows there are new ways to think about more decentralized measures of reputation that are new to this century and ha…

I understand why it seems that nominally objective measures would be better. But I don't think cross-field, non-gameable objective measures of research quality are practically possible I also don't think it's a problem that different groups have different interests, etc. As I say elsewhere, I think that diversity is the solution.

You could be right, but I don't see how it can be known with any confidence until a few approaches are given extended good faith trials. There are anecdotal examples supporting both scenarios and the problem simply seems too unknowable and important not to test drive whatever the top 2 or 3 approaches end up being.

>I also don't think it's a problem that different groups have different interests, etc.

I don't see how you can disagree that cooperation of community to try something different is not a major hurdle.

How many years has it been since important issues in the academic process were widely known? How much success in adoption has there been to date, regarding any fundamental changes?

It seems on its face to be crucial.

Re: Hundreds of extreme self-citing scientists revealed in new database

#210
post #104
post #76

Earlier quoted context omitted.

Aka eigenfactor

My understanding is that eigenfactor rates journals, not individual papers, so if somehow you get low-quality (whatever you want that to mean) papers into nature it has no independent way to realize that your specific paper is low quality. Also eigenfactor is biased towards favoring larger journals, which is not obviously a good thing. It would honestly be really cool if someone did page rank for individual papers. I…

Oh good grief you’re right. This is doubly sad because using an ensemble metric for per-author eigenfactor seems like it would be tractable.

Carl Bergstrom is a smart guy so I suppose the practical implementation of the above must have some wrinkles, but with enough brute force it seems tractable. What I despise more than anything is the gaming that takes place for “impact factor”.

I do OK by standard metrics but would very much like to know where I stand by less easily gamed metrics of influence.

Post reply on HN