GotCloud

From Genome Analysis Wiki
Revision as of 11:34, 18 March 2013 by Mktrost (talk | contribs)
Jump to navigationJump to search

Genomes on the Cloud (GotCloud)

To handle the increasing volume of next generation sequencing and genotyping data, we created and developed software pipelines called Genomes on the Cloud (GotCloud).

GotCloud contains Mapping & Variant Calling Pipelines.

Key Features:

  • Connects sequence analysis tools together in automated pipeline
    • Alignment, quality control, variant calling
  • Robust against unexpected system failure using GNU make
    • easy restart after failure
  • Massively parallel, can run hundreds of jobs
    • Splits large jobs into many pieces
    • Simplifies running on clusters
  • Scalable to tens of thousands of samples
  • Easy to use - Automates series of configurable steps so the user doesn't have to understand/configure/know the many tools required to create high quality results
  • Available on Amazon Web Services (AWS) Elastic Compute Cloud (EC2)
  • Run on local machines/clusters
  • Available via Debian Packages

The pipelines in GotCloud have been used for our sequence data processing and are now publicly available.


Join GotCloud mailing list

Please join in the GotCloud Google Group to ask / discuss / comment about these pipelines.

Currently the "join" button appears to be missing. Click "NEW TOPIC", then select "Join this group". You can then cancel the message post (or post a message).

You can also email Mary Kate Wing (mktrost@umich.edu).


Detailed Background Information

  • Why use GotCloud?
    • Many tools required to create high quality

GotCloudDiagram.png


AWS

The following describes the use of this software with the Amazon Web Services (https://aws.amazon.com/), but you can just as easily use the pipelines on your own machine(s) by just installing them.

Latest Documentation at Tutorial: GotCloud


Setup

You may run the GotCloud software in several modes:

  • On your own hardware running Ubuntu or Redhat/CentOS. See the instructions about installing the software below.
  • On any EC2 instance that uses Ubuntu or Redhat/CentOS distribution. You can install the software as described below, or create a volume using our snapshot (see Amazon Snapshot).
  • On an EC2 cluster instance created by StarCluster. You can install the software as described below, or create a volume using our snapshot (see Amazon Snapshot).

Details for the Choices of Your Install

Install GotCloud Software

Install Resource Files

Resources / Cost

Configure


Running GotCloud Software

Tutorial: GotCloud

Development Notes