Data PUBLISHING: PEER REVIEW, SHARED STANDARDS AND COLLABORATION Rebecca Lawrence, PhD Publisher, F1000 Research rebecca.lawrencef1000.com http:f1000research.com What are we going to talk about About F1000 About F1000 Research
"Data PUBLISHING: PEER REVIEW, SHARED STANDARDS AND" is the property of its rightful owner. Permission is granted to
download and print the materials on this website for personal, non-commercial use only, and to display it
on your personal computer provided you do not modify the materials and that you retain all copyright
notices contained in the materials. By downloading content from our website, you accept the terms of this
agreement.
Presentation Transcript
01
Data PUBLISHING:
PEER REVIEW, SHARED STANDARDSAND COLLABORATION Rebecca Lawrence, PhD
Publisher, F1000 Research
What are we going to talk about About F1000
About F1000 Research
Publication process
Importance of collaboration
Collaborative initiatives and F1000 Research
Challenges of data peer review
Approaches to data peer review
Challenges of post-publication peer review
Summary<br>
03
about F1000 Core service
From the founders of BioMed Central and Current Opinions journals
Post-publication peer review
Faculty of 10,000 experts
Faculty identify and evaluate the most important articles in biology and medicine
1,500 new evaluations per month; >120,000 total so far F1000.com NEW!
F1000 Posters
F1000 Research (F1000R): Plans announced end-Jan; launch later this year<br>
04
F1000R: What are we trying to achieve Alternative to current scholarly publishing approaches tackling 4 problems:
Speed
Immediate publication
Peer review
Open, post-publication peer review
Dissemination of findings
Wide variety of formats
Sharing of primary data
Sharing, publication and refereeing of datasets
Other key features:
‘Gold’ Open Access
Creative Commons CC-BY licences as default
Large (100+), very senior Advisory Panel (e.g. Sir Tim Hunt, Pippa Marrack, Steven Hyman, Alan Schechter, Janet Thornton)<br>
05
F1000R: our publishing process Traditional journal<br>
06
f1000R: Author incentive to make data usable<br>
07
Data publication: many outstanding issues Numerous outstanding issues need to be addressed
Providing benefits even if someone else makes an important discovery from the data – data co-authorship
Effort and time required to sort out the data and dig out the necessary metadata
Lack of formal recognition of data as a valuable output
Technical issues – formats, interoperability, mining tools
Where to store the data, how much to store, and for how long
Data publication: importance of collaboration Some progress has been made:
Growing recognition of the value of data sharing/publication from all stakeholders
Each stakeholder group have made their own advancements
But not going to solve these issues working alone: need to look at the whole ecosystem
Key areas that particularly require stakeholder collaboration (incl cross-publisher):
Workflows involved in the data publication process:
Cross-linking between journals and data repositories
Minimise replication of effort by authors
Format issues
Data repository accreditation
Peer review of datasets<br>
09
Collaborative initiatives and the f1000R Data article F1000R working with other publishers (STM Data Group planned), and many members of all the stakeholder groups on all aspects of the data article.
Datasets
Issues of common/mineable formats (DCXL)
Deposit in relevant subject repositories where possible (BioDBCore)
Otherwise in a stable general data host (Dryad, FigShare, institutional data repository if permanent e.g. Oxford DataBank)
What counts as an ‘approved repository’; what level of permanency guarantees are necessary?
Protocol information
Enough for reuse
Ultimate aim is computer mineable
MIBBI standards too extreme but need some structure
Collaborating on ISA framework development and workflow tools with key groups at Oxford and Harvard Universities<br>
10
F1000r: simplifying and Incentivising data sharing Keep it quick and simple!
Minimal effort
Maximal reuse of experimental and institutional metadata capture
Smooth workflow between article, institutional repositories and data centres
Incentives to share data:
Show view/download statistics – often higher than researchers think
Provide impact measures to show value back to funders, institutions
Encourage data citation in main article references:
Open letter (most major publishers interested in signing)
Scopus/WoK tracking
Agree standard data citation approach<br>
11
Challenges in refereeing data Time required to view, often many, data files (e.g. J Neurosci)
How do you know it is ok?
Without repeating the experiment yourself
Without analysing it yourself<br>
12
Essd journal approach Earth System Science Data journal (Copernicus)
http://www.earth-system-science-data.net/
ESSD peer review ensures that the datasets are:
At least plausible and contain no detectable problems;
Sufficient high quality and their limitations clearly stated;
Well annotated by standard metadata and available from a certified data center/repository;
Customary with regard to their format(s) and/or access protocol, and expected to be useable for the foreseeable future.
Openly accessible (toll free)<br>
13
Pensoft biodiversity data publishing approach Pensoft guidelines for reviewers of data papers
Scientific importance and uniqueness
Data stored in an appropriate repository?
Description of data access?
Complete and uniform recording of the data?
Accurate description of the data?
Use of applicable standards
Possible sources of error appropriately addressed?
Methods to process and analyse the data documented well enough to enable replication?
Data plausible, given the protocols?
All claims substantiated by the underlying data?
Bmc research notes approach Is the question posed original and well defined?
Are the data sound and well controlled?
Is the interpretation well balanced and supported by the data?
Are the methods appropriate and well described; are sufficient details provided to allow others to evaluate and/or replicate the work?
What are the strengths and weaknesses of the methods?<br>
15
F1000r data peer review approach Based on extensive discussion, peer review would focus on:
Is the method used appropriate for the scientific question being asked?
Has enough information been provided to be able to replicate the experiment?
Have appropriate controls been conducted, and the data presented?
Is the data in a useable format/structure?
Are stated data limitations and possible sources of error appropriately described
Does the data ‘look’ ok (optional; e.g. Microarray data)
Our sanity check will pick up:
Format and suitable basic structure adherence
A standard basic protocol structure is adhered to
Data stored in the most appropriate and stable location
Ultimate referee: reuse!<br>
16
F1000R: a two-stage peer review process FIRST: Rapid ‘seems ok’ stamp
SECOND: Subsequent referee comments
All ‘open’
Focus is on whether the work is scientifically sound, not on novelty/interest etc
Encourage author–referee discussion
Encourage author revision (versioning)
Clearly labelled; separate to user commenting<br>
17
challenges of post-publication Refereeing Referee incentives; author revision incentives
Clarity on referee status at any one time
Knowledge of referee status away from the site: CrossMark
Management of several versions; what and how to cite
Simple universally recognisable system for overall referee status:
Approved
Not approved<br>
18
F1000r: Referee status display BETA<br>
19
summary There is now a general consensus that sharing and publishing data is good
Each stakeholder group has made some steps forward
We now need to work together offer some real publishing options to expose data
We need to keep it simple
Develop tools to minimise additional effort required by the research
Develop some common approaches to minimise confusion and wasted effort on applying different formats for each publisher
Build in referee incentives to conduct data peer review
Develop variety of metrics to show the value of submitting your data for peer review; and recognition by those that count