Help Login Create account

Data released on March 27, 2018

Supporting data for "Tracking the NGS revolution: managing life science research on shared high-performance computing clusters"

Dahlö, M; Scofield, D; Schaal, W; Spjuth, O (2018): Supporting data for "Tracking the NGS revolution: managing life science research on shared high-performance computing clusters" GigaScience Database. http://dx.doi.org/10.5524/100421 RIS BibTeX Text

Next-Generation Sequencing (NGS) has transformed the life sciences and many research groups are newly dependent upon computer clusters to store and analyse large datasets. This creates challenges for e-infrastructures accustomed to hosting computationally mature research in other sciences. Using data gathered from our own clusters at UPPMAX computing centre at Uppsala University, Sweden, where core hours usage by ∼800 NGS and ∼200 non-NGS projects is now similar, we compare and contrast the growth, administrative burden and cluster usage of NGS projects with projects from other sciences.
The number of NGS projects has grown rapidly since 2010, with growth driven by entry of new research groups. Storage used by NGS projects has grown more rapidly since 2013 and is now limited by disk capacity. NGS users submit nearly twice as many support tickets per user, and 11 more tools are installed each month for NGS than non-NGS projects. We develop usage and efficiency metrics and show that compute jobs in NGS projects use more RAM than in non-NGS projects, are more variable in core usage, and rarely span multiple nodes. NGS jobs use booked resources less efficiently for a variety of reasons. Active monitoring can improve this somewhat.
Hosting NGS projects imposes a large administrative burden at UPPMAX, due to large numbers of inexperienced users and diverse and rapidly evolving research areas. We give a set of recommendations for e-infrastructures hosting NGS research projects. We provide anonymised versions of our storage, job and efficiency databases.

Contact Submitter

Keywords:

bioinformatics resource usage efficiency efficiency metrics storage next generation sequencing high performance computing e-infrastructures bioinformatics 

Software

http://gigadb.org/images/data/cropped/100421.jpg

Files: (FTP site) Table Settings

Columns:

File Description
Sample ID
Data Type
File Format
Size
Release Date
Download Link
File Attributes

File NameSample IDData TypeFile FormatSizeRelease Date 
Tabular DataTEXT0.21 KB2018-03-02
MD5sumTEXT0.33 KB2018-03-02
ReadmeTEXT1.64 KB2018-03-02
ReadmeTEXT2.33 KB2018-03-02
ReadmeTEXT4.43 KB2018-03-02
scriptR4.66 KB2018-03-02
scriptR6.88 KB2018-03-02
scriptR9.78 KB2018-03-02
scriptR9.85 KB2018-03-02
scriptR10.27 KB2018-03-02
Displaying 1-10 of 18 File(s).

History:

+

Other datasets you might like: