The continued decrease in sequencing costs has led to an abundance of high-throughput data representing an increasing diversity of experimental conditions. These changes have been coupled with the adoption of cloud technologies and interoperability standards to share and analyze large primary and secondary data files. While 10 years ago analysis of hundreds or thousands of genomics samples was only practical at institutions with large local computational resources, these experiments can now be routinely performed by anyone with access to the Internet. In this tutorial, we use the Seven Bridges Cancer Genomics Cloud (CGC) to analyze RNA sequencing data from the NIH Cancer Research Data Commons (CRDC). This tutorial demonstrates how to bring a new computational algorithm to the platform, combine it with an existing workflow, and execute an analysis on the cloud. We highlight best practices for designing command line tools, Docker containers, and CWL descriptions to enable massively parallelized and reproducible biomedical computation with cloud resources. The CGC’s support for diverse analysis techniques and user-friendly interface simplifies the complex process of handling large datasets while promoting collaboration across disciplines.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Building Portable and Reproducible Cancer Informatics Workflows for Scalable Data Analysis: An RNA Sequencing Tutorial

  • Rowan F. Beck,
  • Zelia F. Worman,
  • Gaurav Kaushik,
  • Brandi N. Davis-Dusenbery

摘要

The continued decrease in sequencing costs has led to an abundance of high-throughput data representing an increasing diversity of experimental conditions. These changes have been coupled with the adoption of cloud technologies and interoperability standards to share and analyze large primary and secondary data files. While 10 years ago analysis of hundreds or thousands of genomics samples was only practical at institutions with large local computational resources, these experiments can now be routinely performed by anyone with access to the Internet. In this tutorial, we use the Seven Bridges Cancer Genomics Cloud (CGC) to analyze RNA sequencing data from the NIH Cancer Research Data Commons (CRDC). This tutorial demonstrates how to bring a new computational algorithm to the platform, combine it with an existing workflow, and execute an analysis on the cloud. We highlight best practices for designing command line tools, Docker containers, and CWL descriptions to enable massively parallelized and reproducible biomedical computation with cloud resources. The CGC’s support for diverse analysis techniques and user-friendly interface simplifies the complex process of handling large datasets while promoting collaboration across disciplines.