Showing posts with label metrics. Show all posts
Showing posts with label metrics. Show all posts

Saturday, 16 January 2010

Medical Hypotheses row resurfaces

Last year, publishers Elsevier got into trouble with HIV-AIDS researchers, after Medical Hypotheses (an Elsevier journal) published two papers on the subject of AIDS: one by Peter Duesberg claiming that the AIDS epidemic in South Africa was overhyped, and another by Marco Ruggiero suggesting that the Italian health ministry did not believe that HIV was the sole cause of AIDS (blog posts at Bad Science and Respectful Insolence). The papers were deeply flawed, and were retracted by Elsevier pending an investigation into how they were published. The story has now resurfaced in the Times Higher Education Supplement (THES), because:
Prominent Aids researchers contacted Elsevier to object to the article and wrote to the US National Library of Medicine requesting that Medical Hypotheses be removed from the Medline citation database - an act that would exclude it from the mainstream scientific-communication network.
Elsevier have now convened an expert panel to decide on the future of Medical Hypotheses, with conclusions due by the end of 2010.

In fact, there is no great mystery as to how these flawed papers came to be published. Medical Hypotheses is not peer reviewed: instead, decisions on publication are taken solely by the journal's editor, Prof Bruce Charlton. Articles are often accepted within days, or even hours, of being submitted, suggesting there is little or no quality control on what gets published. Prof Charlton defends this process on two grounds: firstly, that there ought to be some outlet for speculative and bizarre ideas that will not be published by mainstream journals. Secondly, that Medical Hypotheses is a successful and influential journal. Here's what he has to say on the comments following the THES article:
The basic facts are that Medical Hypotheses - www.elsevier.com/locate/mehy - is explicitly and proudly editorially-reviewed (i.e. by me - not peer reviewed); aims to publish radical and revolutionary scientific ideas; and it is objectively a successful journal. It makes a profit, the Thomson ISI Impact Factor is 1.416 (much better than average, and rising), and I know from internal sources that there are half a million papers downloaded per year - which is equivalent download usage to the prestigious Journal of Theoretical Biology. Clearly, in spite or because of our policy to publish bold and sometimes bizarre ideas, Medical Hypotheses plays a significant role in medical science. Fact; not opinion. The editorial advisory board currently includes such respected figures as Nobelist Arvid Carlsson http://en.wikipedia.org/wiki/Arvid_Carlsson; Sir Roy Calne http://en.wikipedia.org/wiki/Roy_Calne; Antonio Damasio http://en.wikipedia.org/wiki/Antonio_Damasio and V.S. Ramachandran http://en.wikipedia.org/wiki/Vilayanur_S._Ramachandran . Past editorial advisors have included Sir Karl Popper and Nobelist Sir James Black. *** There are only two possible legitimate outcomes to the current process. Either: 1. Medical Hypotheses could continue as an influential, profitable and well-known editorially-reviewed journal with a radical mission. Or else: 2. The journal could be closed-down altogether, and the title abolished. But it would obviously not be ethically acceptable to launch a new ‘imposter’ journal - with utterly different editorial aims, procedures and personnel; yet retaining the 34 year established title of Medical Hypotheses.
As I keep saying, the impact factor of a journal tells you nothing about its quality. For example, here are three peer-reviewed pseudojournals that repeatedly publish abject nonsense and pdeudoscience, with their impact factors according to Journal Citation Reports:
  • Homeopathy: 1.041
  • Evidence-based Complementary and Alternative Medicine: 1.954
  • Journal of Alternative and Complementary Medicine: 1.628
The articles in these journals are typically written by quacks, and are cited by other quacks writing in quack journals, giving a high-ish but meaningless impact factor. Perhaps Medical Hypotheses is also highly influential among pseudoscientists?

But the main point here is about radical and controversial hypotheses. I think most people would agree that these have their place in scientific discourse, and there ought to be somewhere to publish them. However, this isn't really what the argument is about. In this case, two fatally flawed papers were published with little or no scrutiny: these papers have potential global health implications. In the case of the Duesberg paper, reviews posted on the Denying AIDS blog show the major problems with the paper. There's a difference between publishing provocative ideas that might inspire new research, and ones that are just demonstrably wrong. While the likes of Peter Duesberg have the right to say what they like, they don't have the right to say it in a MEDLINE-indexed journal. This is not an argument about free speech, it's an argument about the integrity of the scientific literature. There may be a place for journals such as Medical Hypotheses, but there has to be some level of quality control. Otherwise, why should anyone take them seriously?

Wednesday, 24 June 2009

What do bibliometrics actually add to research evaluation?

Firstly, the reason that I haven't posted in an age is that I've been in Norway, interpreting seismic data for the new project I'm working on. Hopefully I can now post a bit more regularly, as I should actually be in Manchester for a few consecutive weeks, for the first time this year.

Regular readers will know that I like to whinge about the increasing use of statistical indicators (bibliometrics) to evaluate research performance. Previously in England, research performance has been evaluated by the Research Assessment Exercise, a cumbersome and involved system based around expert peer review of research. Currently, HEFCE (the body that decides how scarce research funding is allocated to English universities) is looking into replacing this with a cumbersome and involved system based around bibliometrics and "light-touch" peer review. To this end, a pilot exercise using bibliometrics and including 22 universities has been underway. An interim report on the pilot is now available.

Essentially, three approaches have been evaluated:

i) Based on institutional addresses: here papers are assigned to a university based on the addresses of the the authors, as stated in the paper. This would be cheap to do, as it would need no input from the universities.

ii) Based on all papers published by authors. In this approach, all papers written by staff selected for the 2008 RAE were identified. This requires a lot of data to be collected.

iii) Based on selected papers published by authors. Again, this approach used all staff selected for the 2008 RAE, but only used the most cited papers.

For each approach, the exercise was conducted twice: once using the Web Of Science (WoS) database, and once using Scopus. The results were then compared with those from the 2008 RAE.

Well, the results are interesting, if you like this sort of thing. It is clear that the results can be very different from those provided by the RAE, whichever method was used, although the "selected papers" method tends to give the closest results. It is also notable that the two different databases give different results, sometimes radically so; Scopus seems to consistently give higher values than WoS. Workers in some fields complained that they made more use of other databases, such as the arXiv or Google Scholar (it's worth noting that the favoured databases are proprietary, while the arXiv and Google Scholar are publically accessible).

In general, the institutions involved in the pilot preferred the "selected papers" method, but it seems that none of the methods produced particularly convincing results. According to the report (paras 66 and 67):

In many disciplines (particularly in medicine, biological and physical sciences and psychology), members reported that the ‘top 6’ model (which looked at the most highly cited papers only) generally produced reasonable results, but with a number of significant discrepancies. In other disciplines (particularly in the social sciences and mathematics) the results were less credible, and in some disciplines (such as health sciences, engineering and computer science) there was a more mixed picture. Members generally reported that the other two models (which looked at ‘all papers’) did not generally produce credible results or provide sufficient differentiation.

One of the questions here is what is meant by "reasonable" or "credible" results? The institutions involved in the pilot seem to assume that the best results are the ones that most closely match those of the RAE. I suspect this is because the large universities that currently receive the lion's share of research funding are not going to support any system that significantly changes the status quo.

The institutions involved in the pilot seem to think that bibliometrics would be most useful when used in conjunction with expert peer review. From the report:

Members discussed whether the benefits of using bibliometrics would outweigh the costs. Some found this difficult to answer given limited knowledge about the costs. Nevertheless there was broad agreement that overall the benefits would outweigh the costs – assuming a selective approach. For institutions this would involve a similar level of burden to the RAE and any additional cost of using bibliometrics would be largely absorbed by internal management within institutions. For panels, some members felt that bibliometrics might involve additional work (for example in resolving differences between panel judgements and citation scores); others felt that they could be used to increase sampling and reduce panels’ workloads.

According to the interim report, the "best" results (i.e. those most closely matching the results of the RAE) were obtained using a methodology that will have a similar administrative burden as the RAE. Even then the results had "significant discrepancies". So, if the aim of the pilot was to get similar results to the RAE with a lesser administrative burden, it seems that the pilot exercise has failed on both counts. So if bibliometrics don't seem to add much to the process, it's worth considering what they might take away. For which, see my previous post...

Friday, 9 January 2009

Does the REF add up to good science?

The RAE (Research Assessment Exercise) results from the 2008 were published back in December. You might have noticed this from the number of university websites that could be found frantically spinning the results. My very own University of Manchester, for example, is claiming that Manchester had broken into the “golden triangle” of UK research, that is, Oxford, Cambridge and institutions based in London. It seems that depending on the measure you pick, we’re anywhere between third and sixth place in the UK. Clearly these are excellent results, but whether we’re really up there with the Oxfords, Cambridges, Imperials and UCLs of the world I’m not sure.

In any case, that was the last ever RAE. It has been a fairly cumbersome process, involving expert peer review of the research contribution of research institutions, that has been a real burden on the academics who have had to administer it. I’m sure there are few who will mourn its passing. Now the world of English academia is waiting, like so many rats in an experimental maze, to find out what will replace the RAE. The replacement will be a thing called the Research Excellence Framework, or REF, and at this stage exactly what it will involve is fairly sketchy. However, it will be based on the use of bibliometrics (statistical indicators that are usually based on how much published work is cited in other publications) and “light-touch peer review”.

What kind of bibliometric indicators are we talking about? Last year HEFCE (the Higher Education Funding Council for England, the body that evaluates research and decides who gets scarce research funding) published a “Scoping study on the use of bibliometric analysis to measure the quality of research in UK higher education institutions” produced by the Centre for Science and Technology Studies at the University of Leiden, Netherlands. I’ve spent a fair amount of time reading through this, and in some ways I was encouraged. It’s clear that some thought has gone into creating bibliometric indicators that are as sensible as possible: I was dreading a crude approach based around impact factors, which have already done so much damage to the pursuit of good science. The authors of the “scoping study” came up with an “internationally standardised impact indicator”: I will abbreviate this as ISII for concision. The ISII takes the average number of citations for publications for the academic unit you are interested in (this might be a research group, an academic department or an entire university), and divides it by a weighted, field-specific international reference level. The reference level is calculated by taking the average number of citations for all publications in a specific field: if the publication falls under more than one field (as many will in practice), the reference level can be calculated as a weighted average of the number of citations generated by publications in all the fields in question. So, if the ISII for your research group comes out as 1, you’re average, if above 1, better than the average, and if below 1, worse than the average. The authors of the scoping study say that they regard the ISII as being “the most appropriate research performance indicator”, and suggest that a value of >1.5 indicates a scientifically strong institution. They also suggest a threshold of 3.0 to identify research excellence. It seems that the HEFCE is expecting to adopt the ISII as the main research performance indicator, according to their FAQs, where they say “We propose to measure the number of citations received by each paper in a defined period, relative to worldwide norms. The number of citations received by a paper will be 'normalised' for the particular field in which it was published, for the year in which it was published, and for the type of output”. However, they are still deciding what thresholds they will use to decide which institutions are producing high-quality research.

All well and good. If you insist that bibliometric indicators are necessary, this is probably as good a way as any of generating those data. However, there are some problems here, as well as philosophical difficulties with the entire approach.

Firstly, what is it we are trying to measure? In theory, what HEFCE wants to do is evaluate research quality. But the ISII does not directly measure research quality. Like any indicator based on citation rates, it is measuring the “impact” of the research: how many other researchers published papers that cited the research. It ought to be clear that while this should reflect quality to some degree, there are significant confounding factors. For example, research that is done in a highly active topic is likely to be cited more than research in which fewer groups are working. This does not mean that work in less active topics is of intrinsically lower quality, or even that it is less useful.

Secondly, there is an assumption that the be-all and end-all of scientific research is publication in peer-reviewed journals that are indexed in the Web of Science citation database published by Thomson Scientific. This a proprietary database that lists articles in the journals that it indexes, and also tracks citations. Criteria for journals to be included are not in the public domain (although the scoping report suggests these are picked based on their citation impact, p. 43). A number of journals that I would not consider to be scientifically reputable are included. For example, under the heading of Integrative and Complementary Medicine, the 2007 Journal Citation Reports (a database that compiles bibliometric statistics for journals in the citation database) includes 12 journals, including Evidence Based Complementary and Alternative Medicine (impact factor 2.535!) and the Journal of Alternative and Complementary Medicine (impact factor 1.526). This reinforces the point made above: it would be possible to publish outright quackery in either of these journals, have it cited by other quacks in the quackery that they publish, and get a respectable rating on the ISII. The ISII can’t tell you that this is a vortex of nonsense: it only sees that other authors have cited the work. It is also true that not all journals are included in the citation index: for example, in my own field the Bulletin of Canadian Petroleum Geology fails to make the cut, although it has always published good quality research. Although the authors of the scoping report make clear that it is possible to expand bibliometrics beyond the citation database, this will take much more effort and it seems that HEFCE will not take this route. So we will be relying on a proprietary and opaque database to make decisions on future research funding. A further point is that it is not clear how open access publications will be incorporated in the citation index: in principle there is no reason that this can’t happen, but can we be sure it will?

Thirdly, there is the assumption that research output can only be evaluated in terms of published articles in peer-reviewed journals. I’m not sure that this accurately reflects the actual research output of many scientists. For example, most of us put a lot of effort into presentations at scientific conferences, chapters in books, or government reports that will never make it into a citation database. This has become a problem for things like, in my own field, the special publications of the Geological Society of London. These are volumes that collect recent research on specific topics, and they generally contain excellent research. But they aren’t included in citation databases and they have no impact factor. This has led to a lack of interest in publishing results in these special publications, because they don’t tick the right boxes in terms of publication metrics. This is surely a bad thing. A similar problem occurs with things like government open-file reports. These are not, in general, pieces of world-class, cutting edge research. But that does not mean that they are useless or that they have no value. For example, good regional geological work can allow mineral exploration to be better targeted, benefiting the local economy. Yet that kind of work is ignored in a framework that only considers journal articles: HEFCE says only that “We accept that citation impact provides only a limited reflection of the quality of applied research, or its value to users. We invite proposals for additional indicators that could capture this”. To me, research quality and value cannot be measured by bibliometric indicators. It can only be evaluated by reading the research, understanding its context within the totality of pre-existing research, and understanding how it contributes to new understanding. That is, it can only be evaluated through peer review.

Which brings me to my fourth point; there are some questions about the role of peer review within the REF. HEFCE says that “the scoping study recommends that experts with subject knowledge should be involved in interpreting the data. It does not recommend that primary peer review (reading papers) is needed in order to produce robust indicators that are suitable for the purposes of the REF”. However, I’m not convinced that this accurately summarises what is written in the scoping report, which says “In the application of indicators, no matter how advanced, it remains of the utmost importance to know the limitations of the method and to guard against misuse, exaggerated expectations of non-expert users, and undesired manipulations by scientists themselves…Therefore, as a general principle we state that optimal research evaluation is realised through a combination of metrics and peer review. Metrics, particularly advanced analysis, provides the tools to keep the peer review process objective and transparent. Metrics and peer review both have their strengths and limits. The challenge is to combine the two methodologies in such a way that the strengths of one compensates for the limitations of the other”.

Finally, there is a hint of conflict of interest in the preparation of the scoping report by the Centre for Science and Technological Studies: according to their website, the centre is involved in selling "products" based on its research and development in the area of bibliometric indicators. Their report in favour of bibliometric indicators might allow them to drum up significant business from HEFCE.

At present, the proposals for the REF are at a fairly early stage, but the use of bibliometric indicators seems to be entrenched, and there will be a pilot exercise on bibliometric indicators this year. However, this is based on “expert advice” that consists of a single report from an organisation that makes money by creating bibliometric indicators. While academia in general might welcome the proposals on the grounds that they will be less burdensome than the RAE and give everyone more time to do research, I don’t think many academics will be kidding themselves that the bibliometric indicators involved actually tell us much about research quality and usefullness.

Monday, 4 June 2007

Metrication

On those rare occasions when I actually have a paper to submit for publication, I tend to consider the most appropriate journal for the article I've written. I weigh up factors such as the likely readership of the paper, the readership of the journal, the journal's reputation for rapid peer review and editorial processes, and so on. I have colleagues who say that I ought to take things like the impact factor into account, because publishing in journals with high impact factors is important for my career. The question is, why?

The impact factor is calculated using a publication database by Thomson Scientific, and published on the ISI Web of Knowledge. The database is proprietary, but my employer subscribes to it. The impact factor basically works by counting the number of citations in the year to articles published in a particular journal over the last two years, and dividing by the total number of articles published over the last two years. So the 2006 numbers, which will be published at the end of this month, are derived by counting the number of citations from articles published in 2006 to articles published in the journal in question in 2004 and 2005, and dividing by the total number of articles published in the journal in question in 2004 and 2005. The impact factor applies to the journal as a whole, and not the individual papers published in it.

Why do I hate the impact factor so? There are several reasons. Firstly, it is statistically dubious. Journals with high impact factors get most of their citations from a relatively small number of highly cited papers. For example, David Colquhoun writes that in Nature in 1999 the most cited 16% of papers accounted for 50% of citations (Nature 423, p. 479). In other words, the citation rate of an individual paper is uncorrelated to the impact factor of a journal. This is the most obvious statistical drawback. Several other sources of bias are described in Seglen (1997; British Medical Journal 314, p.497). Eugene Garfield, who invented the impact factor, also agrees that the impact factor should not be used to evaluate individuals (Garfield 1998; Der Unfallchirurg 48, p. 413).

Another 'metric' that has been suggested to evaluate the contribution of scientists is the h-index. In this scheme, an author will have an h-index of h such that they have published h papers that have each been cited at least h times. It has been pointed out that Einstein, had he died in early 1906, would have an h of only 4 or 5, despite the revolutionary nature of the work he had published before that date. This makes me slightly happier about my own h-index of 0, being early in my career and having published two papers that have not yet been around long enough to be cited.

Really though, statistical arguments about biases in various metrics miss the central point, which is that it is impossible to evaluate the scientific worth of an article without reading and understanding it. This can only be done well by people who have expertise in the field of study. In other words, it can only be done well by peer review.

In the UK, academic researchers are evaluated through the Research Assessment Exercise (RAE), which has traditionally been based on peer review. In the upcoming RAE in 2008, there will be a 'shadow' metrics-based exercise running alongside the traditional peer-review based process. In RAEs after 2008, metrics will be used as the main measure of the scientific worth of individual researchers. There have been many criticisms of the traditional RAE. David Colquhoun has written "All of us who do research (rather than talk about it) know the disastrous effects that the Research Assessment Exercise has had on research in the United Kingdom: short-termism, intellectual shallowness, guest authorships and even dishonesty" (Nature 446, p. 373). The situation is hardly going to be improved by relying on a metrics-based approach, as authors inevitably play the system in order to inflate their rankings and progress their careers.

What is perhaps most disappointing is that in general scientists themselves seem to have failed to critically examine metrics such as the impact factor. There is enough material out there in the public domain (the sources cited here are only a sample) for anyone to understand the problems, if they're interested in finding out.