<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Srashkin</id>
	<title>Genome Analysis Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Srashkin"/>
	<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/wiki/Special:Contributions/Srashkin"/>
	<updated>2026-09-25T08:03:32Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.43.1</generator>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Sara_Rashkin&amp;diff=7552</id>
		<title>Sara Rashkin</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Sara_Rashkin&amp;diff=7552"/>
		<updated>2013-06-27T16:42:06Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Sara is a PhD student working with [[Goncalo Abecasis]] in developing methods for copy number variation detection and is currently working on identifying optimal strategies for identifying disease associated singletons.  She earned a BA in statistics from Northwestern University, where she graduated with departmental honors, and an MS in biostatistics from the University of Michigan.&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Sara_Rashkin&amp;diff=7551</id>
		<title>Sara Rashkin</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Sara_Rashkin&amp;diff=7551"/>
		<updated>2013-06-27T16:41:03Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: Created page with &amp;#039;Sara is a PhD student working with Goncalo Abecasis in developing methods for copy number variation detection as well as identifying optimal strategies for identifying diseas…&amp;#039;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Sara is a PhD student working with [[Goncalo Abecasis]] in developing methods for copy number variation detection as well as identifying optimal strategies for identifying disease associated singletons.  She earned a BA in statistics from Northwestern University, where she graduated with departmental honors, and an MS in biostatistics from the University of Michigan.&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Abecasis_Lab&amp;diff=7550</id>
		<title>Abecasis Lab</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Abecasis_Lab&amp;diff=7550"/>
		<updated>2013-06-27T16:34:40Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: /* Graduate Students */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Image:Abecasis_group_photo_cropped.jpg|900px|center|Group Photo 2013]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--[[Image:2009.08_Group_Retreat_Photo.jpg|400px|center|Group Photo]]--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Mission Statement ==&lt;br /&gt;
&lt;br /&gt;
We are developing and applying computational and statistical tools to further understanding of complex human diseases, such as cardiovascular disease and diabetes.&lt;br /&gt;
&lt;br /&gt;
== Leadership ==&lt;br /&gt;
&lt;br /&gt;
[[Goncalo Abecasis]] is currently the Felix Moore Collegiate Professor of Biostatistics at the University of Michigan School of Public Health.&lt;br /&gt;
&lt;br /&gt;
== Current Members ==&lt;br /&gt;
&lt;br /&gt;
=== Research Fellows ===&lt;br /&gt;
&lt;br /&gt;
* Goo Jun&lt;br /&gt;
* Christian Fuchsberger&lt;br /&gt;
* Alex Tsoi&lt;br /&gt;
* [[Dajiang Liu]]&lt;br /&gt;
* [[Lars Fritsche]]&lt;br /&gt;
* [[Scott Vrieze]]&lt;br /&gt;
&lt;br /&gt;
=== International Visitors ===&lt;br /&gt;
&lt;br /&gt;
* Andrea Maschio&lt;br /&gt;
* Giorgio Pistis&lt;br /&gt;
* Eleanora Porcu&lt;br /&gt;
&lt;br /&gt;
=== Graduate Students ===&lt;br /&gt;
&lt;br /&gt;
* Su Chu&lt;br /&gt;
&lt;br /&gt;
* Sayantan Das&lt;br /&gt;
&lt;br /&gt;
* Shuang Feng&lt;br /&gt;
&lt;br /&gt;
* Dan Hovelson&lt;br /&gt;
&lt;br /&gt;
* [[Alan Kwong]]&lt;br /&gt;
&lt;br /&gt;
* Ben Lerch&lt;br /&gt;
&lt;br /&gt;
* [[Sara Rashkin]]&lt;br /&gt;
&lt;br /&gt;
* Sebanti Sengupta&lt;br /&gt;
&lt;br /&gt;
* Vivian Wang&lt;br /&gt;
&lt;br /&gt;
* Xiaowei Zhan&lt;br /&gt;
&lt;br /&gt;
* Tingting Zhou&lt;br /&gt;
&lt;br /&gt;
=== Staff ===&lt;br /&gt;
&lt;br /&gt;
* Laura Baker&lt;br /&gt;
&lt;br /&gt;
* Tom Blackwell&lt;br /&gt;
&lt;br /&gt;
* [[Sean Caron]]&lt;br /&gt;
&lt;br /&gt;
* [[Jennifer Bragg-Gresham]]&lt;br /&gt;
&lt;br /&gt;
* Kevin Li&lt;br /&gt;
&lt;br /&gt;
* [[Mary Kate Wing]]&lt;br /&gt;
&lt;br /&gt;
== Alumni ==&lt;br /&gt;
&lt;br /&gt;
=== Former Research Faculty ===&lt;br /&gt;
&lt;br /&gt;
Hyun Min Kang (&#039;&#039;graduated in 2011&#039;&#039;), now Assistant Professor at the [http://www.sph.umich.edu/biostat/ University of Michigan School of Public Health, Department of Biostatistics].&lt;br /&gt;
&lt;br /&gt;
=== Former Research Fellows ===&lt;br /&gt;
&lt;br /&gt;
Weimin Chen (graduated 2007), now Assistant Professor at the [http://people.virginia.edu/~wc9c/ Department of Public Health Sciences &amp;amp; Center for Public Health Genomics, University of Virginia]&lt;br /&gt;
&lt;br /&gt;
Bingshan Li (graudated 2011), now Assistant Professor at the [https://medschool.vanderbilt.edu/cqs/people/Bingshan/Li/cqs-faculty-members Center for Quantitative Sciences, Vanderbilt University]&lt;br /&gt;
&lt;br /&gt;
Serena Sanna (graduated 2007), now an investigator at the [http://www.serenasanna.com/ Istituto di Neurogenetica e Neurofarmacologia in Sardinia, Italy]&lt;br /&gt;
&lt;br /&gt;
Paul Scheet (graduated 2008), now Assistant Professor at [http://faculty.mdanderson.org/Paul_Scheet/Default.asp?SNID=221605974 Department of Epidemiology, University of Texas MD Anderson Cancer Center]&lt;br /&gt;
&lt;br /&gt;
Carlo Sidore (graduated 2012), now an investigator at the [http://www.serenasanna.com/ Istituto di Neurogenetica e Neurofarmacologia in Sardinia, Italy]&lt;br /&gt;
&lt;br /&gt;
William Stewart (graduated 2008), now Assistant Professor at [http://www.mathmed.org/#William_Stewart Battelle Center for Computational Medicine, Departments of Statistics and Pediatrics, National Children&#039;s Hospital and Ohio State University]&lt;br /&gt;
&lt;br /&gt;
=== Former Doctoral Students ===&lt;br /&gt;
&lt;br /&gt;
Wei Chen (graduated 2011), now Assistant Professor  at the [http://www.chp.edu/CHP/Chen%2C+Wei%2C+PhD Department of Pediatrics, University of Pittsburgh Medical Center]&lt;br /&gt;
&lt;br /&gt;
Jun Ding (graduate 2010), now Staff Scientist / Facility Head at the [http://www.grc.nia.nih.gov/branches/lg/lg.htm Laboratory of Genetics, National Institute on Aging (NIH)].&lt;br /&gt;
&lt;br /&gt;
Yun Li (graduated 2009), now Assistant Professor at the [http://www.sph.unc.edu/?option=com_profiles&amp;amp;Itemid=6138&amp;amp;profileAction=ProfDetail&amp;amp;pid=708777879 Department of Biostatistics, University of North Carolina].&lt;br /&gt;
&lt;br /&gt;
Youna Hu (graduated 2012), now a Research Fellow [http://cteg.berkeley.edu/members/hu.html working with Rasmus Nielsen at Berkeley]&lt;br /&gt;
&lt;br /&gt;
Mingyao Li (graduated 2005), now Associate Professor at the [http://www.cceb.upenn.edu/faculty/index.php?id=159 Department of Biostatistics and Epidemiology, University of Pennsylvania]&lt;br /&gt;
&lt;br /&gt;
Liming Liang (graduated 2009), now Assistant Professor at the [http://www.hsph.harvard.edu/faculty/liming-liang/ Departments of Biostatistics and Epidemiology, Harvard University]&lt;br /&gt;
&lt;br /&gt;
Tasha Fingerlin (graduated 2003), now Associate Professor at the [http://www.ucdenver.edu/academics/colleges/PublicHealth/departments/Epidemiology/About/Faculty/Pages/FingerlinT.aspx Section of Epidemiology and Community Health, University of Colorado Health Sciences Center]&lt;br /&gt;
&lt;br /&gt;
Andrew Skol (graduated 2006), now Assistant Professor at the [http://med-www02.bsd.uchicago.edu/339/FacultyPro/faculty_profile.aspx?empl_id=10164 Section of Genetic Medicine, University of Chicago]&lt;br /&gt;
&lt;br /&gt;
Jin Zhen (graduated 2009), now working in the Pharmaceutical Industry.&lt;br /&gt;
&lt;br /&gt;
=== Former Masters Students ===&lt;br /&gt;
&lt;br /&gt;
Melinda Curran&lt;br /&gt;
&lt;br /&gt;
Vesela Gateva&lt;br /&gt;
&lt;br /&gt;
Xijing Han&lt;br /&gt;
&lt;br /&gt;
Elizabeth Jewell&lt;br /&gt;
&lt;br /&gt;
Yanming Li&lt;br /&gt;
&lt;br /&gt;
Heather Munro&lt;br /&gt;
&lt;br /&gt;
Theresa Scott (nee Daigneault)&lt;br /&gt;
&lt;br /&gt;
Matthew Snyder&lt;br /&gt;
&lt;br /&gt;
Yuan Wei&lt;br /&gt;
&lt;br /&gt;
Abigail Woodroffe&lt;br /&gt;
&lt;br /&gt;
Zaojun Ye&lt;br /&gt;
&lt;br /&gt;
Matthew Zawitowski&lt;br /&gt;
&lt;br /&gt;
Anita Yu Zhao&lt;br /&gt;
&lt;br /&gt;
=== Visitors ===&lt;br /&gt;
&lt;br /&gt;
Toshiko Tanakato&lt;br /&gt;
&lt;br /&gt;
== Really Useful Stuff ==&lt;br /&gt;
&lt;br /&gt;
* [[Abecasis Group Awards]]&lt;br /&gt;
* [https://calendars.office.microsoft.com/pubcalstorage/m3n2kr0z1470909/Goncalo_Abecasis_Calendar(1).ics Goncalo&#039;s Calendar]&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=643</id>
		<title>Base Caller Summaries</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=643"/>
		<updated>2010-03-12T21:04:23Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: /* References */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Standard Illumina Base Caller (Bustard)==&lt;br /&gt;
&lt;br /&gt;
===Sequencing-by-Synthesis (SBS)===&lt;br /&gt;
&lt;br /&gt;
*DNA sample obtained, containing many copies of same sequences and randomly fragmented&lt;br /&gt;
*Single-stranded DNA fragments attached to slide and amplified so there is a cluster of each fragment&lt;br /&gt;
*DNA polymerase and 4 terminal bases (with distinct fluorescent markers) added&lt;br /&gt;
*Clusters excited by lasers and photos taken in optimal wavelengths for 4 fluorophores&lt;br /&gt;
*Fluorophores and terminators removed and process repeated for L cycles&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Corrects for imperfect repositioning of camera and aberrations of lens by aligning images to reference from original cycle&lt;br /&gt;
*Signal for each cluster characterized as time series data of fluorescence intensities and noise&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Converts fluorescence signals into actual sequence data with quality scores&lt;br /&gt;
*Takes intensities of four channels for every cluster in each cycle and determines concentration of each base&lt;br /&gt;
*Renormalizes concentrations by multiplying by ratio of average concentrations in first cycle and current cycle&lt;br /&gt;
*Uses Markov model to determine transition matrix modeling probability of phasing (no new base synthesized), prephasing (two new bases synthesized), and normal incorporation&lt;br /&gt;
*Uses transition matrix and observed concentrations of each base to determine concentrations in absence of phasing and reports these as base calls&lt;br /&gt;
**Assumes crosstalk matrix constant for a given sequencing run and that phasing affects all nucleotides in the same way&lt;br /&gt;
&lt;br /&gt;
===General Noise Factors===&lt;br /&gt;
*Phasing&lt;br /&gt;
**Failures in nucleotide incorporation or block removal or incorporation of more than one nucleotide in a particular cycle&lt;br /&gt;
*Fading&lt;br /&gt;
**Decay in fluorescent signal intensity with each cycle&lt;br /&gt;
**Likely attributable to material loss during sequencing&lt;br /&gt;
*Crosstalk&lt;br /&gt;
**C channel illumination overlaps with A: a C label fluoresces in A channel (similarly G and T overlap)&lt;br /&gt;
**Likely caused by overlap in dye emission frequencies&lt;br /&gt;
*T Accumulation&lt;br /&gt;
**The fluorophores used for thymine are not always removed properly after each iteration&lt;br /&gt;
**Intensity of T signal increases across sequencing run&lt;br /&gt;
&lt;br /&gt;
==Alta-Cyclic==&lt;br /&gt;
===Training Stage===&lt;br /&gt;
*Learns run-specific noise patterns according to model and finds optimized solution reducing affect of noise sources using a Support Vector Machine (SVM)&lt;br /&gt;
*Half of training set used for cross-validation&lt;br /&gt;
&lt;br /&gt;
===Base Calling Stage===&lt;br /&gt;
*Reports all sequences from run with optimized parameters&lt;br /&gt;
&lt;br /&gt;
===Differences from Standard Illumina Base Caller===&lt;br /&gt;
*Calling parameters optimized empirically and tested to enhance accuracy of each run&lt;br /&gt;
*Calculates phasing parameters based on parametric model&lt;br /&gt;
*Dynamically tracks changes in crosstalk, which disrupt signals in later cycles&lt;br /&gt;
&lt;br /&gt;
==Probabilistic Base Calling==&lt;br /&gt;
*Produces an alternative probabilistic base calling method based on the fluorescence intensity quantifications that uses:&lt;br /&gt;
**Extended IUPAC alphabet to code ambiguous bases &lt;br /&gt;
**Information criterion to control length of trustable reads&lt;br /&gt;
*Reduced systematic bias by addressing:&lt;br /&gt;
**Crosstalk&lt;br /&gt;
**Dephasing&lt;br /&gt;
**Optical effect that tiles in center of image appear brighter corrected by fitting a 2D loess model to intensities and subtracting difference between fit and median intensities&lt;br /&gt;
*Measure level of uncertainty in base calling by entropy (uncertainty in determination of correct kth base)&lt;br /&gt;
*Does not consider fine-tuning image analysis&lt;br /&gt;
&lt;br /&gt;
==BayesCall==&lt;br /&gt;
*Model-based approach to base calling&lt;br /&gt;
*Main goal is to model sequencing process by taking stochasticity into account and by explicitly modeling how errors may arise&lt;br /&gt;
*Obtain base calls by maximizing posterior distribution of sequences given observed data and assuming a uniform prior on sequences&lt;br /&gt;
&lt;br /&gt;
==Swift==&lt;br /&gt;
Performs both image analysis and base calling&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Background subtraction – minimal pixel value within a window around each pixel subtracted from central pixel’s value&lt;br /&gt;
*Image correlation – alignment of images to reference cycle&lt;br /&gt;
*Object identification and intensity extraction&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Corrects for crosstalk by performing linear regression on crosstalk plots and use slope to derive correction matrix, performed iteratively until slope is zero&lt;br /&gt;
*Phasing correction by ranking clusters by chastity (the ratio of the highest intensity to the sum of the top two intensities) - use top 400 clusters to estimate phasing and apply it as a correction &lt;br /&gt;
*After correction, base with maximum intensity chosen as called base&lt;br /&gt;
&lt;br /&gt;
==Ibis==&lt;br /&gt;
&lt;br /&gt;
===Method===&lt;br /&gt;
*Estimate sequencing chemistry model as a parameter directly from data using statistical learning&lt;br /&gt;
*Training set from Bustard output using raw cluster intensities&lt;br /&gt;
*Used a base caller with SVM classifiers for each cycle that have intensity values of the current cycle as well as the previous and following cycles (if they exist)&lt;br /&gt;
*Data set created by aligning raw reads with mismatches for a fraction of the tiles to a reference sequence&lt;br /&gt;
**Half of this set used as a training set and the other half as a test set used to check results of training&lt;br /&gt;
*Estimate parameters for calculating a quality score given class assignment and distances to the classification/decision boundary from SVM&lt;br /&gt;
&lt;br /&gt;
===Comparison===&lt;br /&gt;
*Unlike AltaCyclic, includes base-specific phasing parameters so can correct raw intensities for T accumulation&lt;br /&gt;
*Does not call an &#039;N&#039; character for poor quality bases&lt;br /&gt;
*Process unique as causes of sequencing error not modeled separately&lt;br /&gt;
**Consider causes together by using neighboring signals in statistical learning procedure&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&lt;br /&gt;
Erlich, Y., Mitra, P.P., delaBastide, M., McCombie, W.R., Hannon, G.J. (2008) Alta-Cyclic: A self-optimizing base caller for next-generation sequencing. &#039;&#039;Nature Methods&#039;&#039; &#039;&#039;&#039;5&#039;&#039;&#039;:679-682 &lt;br /&gt;
&lt;br /&gt;
Kao, W.-C., Stevens, K., Song, Y.S. (2009) BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing. &#039;&#039;Genome Research&#039;&#039; &#039;&#039;&#039;19&#039;&#039;&#039;:1884-1895 &lt;br /&gt;
&lt;br /&gt;
Kircher, M., Stenzel, U., Kelso, J.  (2009) Improved base calling for the Illumina Genome Analyzer using machine learning strategies. &#039;&#039;Genome Biol.&#039;&#039; &#039;&#039;&#039;10(8)&#039;&#039;&#039;:Article R83&lt;br /&gt;
&lt;br /&gt;
Rougemont, J., Amzallag, A., Iseli, C., Farinelli, L., Xenarios, I., Naef, F. (2008) Probabilistic base calling of Solexa sequencing data. &#039;&#039;BMC Bioinformatics&#039;&#039; &#039;&#039;&#039;9&#039;&#039;&#039;:Article 431 &lt;br /&gt;
&lt;br /&gt;
Whiteford, N., Skelly, T., Curtis, C., Ritchie, M.E., Löhr, A., Zaranek, A.W., Abnizova, I., Brown, C. (2009) Swift: Primary data analysis for the Illumina Solexa sequencing platform. &#039;&#039;Bioinformatics&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:2194-2199&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=642</id>
		<title>Base Caller Summaries</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=642"/>
		<updated>2010-03-12T20:57:06Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: /* Comparison */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Standard Illumina Base Caller (Bustard)==&lt;br /&gt;
&lt;br /&gt;
===Sequencing-by-Synthesis (SBS)===&lt;br /&gt;
&lt;br /&gt;
*DNA sample obtained, containing many copies of same sequences and randomly fragmented&lt;br /&gt;
*Single-stranded DNA fragments attached to slide and amplified so there is a cluster of each fragment&lt;br /&gt;
*DNA polymerase and 4 terminal bases (with distinct fluorescent markers) added&lt;br /&gt;
*Clusters excited by lasers and photos taken in optimal wavelengths for 4 fluorophores&lt;br /&gt;
*Fluorophores and terminators removed and process repeated for L cycles&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Corrects for imperfect repositioning of camera and aberrations of lens by aligning images to reference from original cycle&lt;br /&gt;
*Signal for each cluster characterized as time series data of fluorescence intensities and noise&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Converts fluorescence signals into actual sequence data with quality scores&lt;br /&gt;
*Takes intensities of four channels for every cluster in each cycle and determines concentration of each base&lt;br /&gt;
*Renormalizes concentrations by multiplying by ratio of average concentrations in first cycle and current cycle&lt;br /&gt;
*Uses Markov model to determine transition matrix modeling probability of phasing (no new base synthesized), prephasing (two new bases synthesized), and normal incorporation&lt;br /&gt;
*Uses transition matrix and observed concentrations of each base to determine concentrations in absence of phasing and reports these as base calls&lt;br /&gt;
**Assumes crosstalk matrix constant for a given sequencing run and that phasing affects all nucleotides in the same way&lt;br /&gt;
&lt;br /&gt;
===General Noise Factors===&lt;br /&gt;
*Phasing&lt;br /&gt;
**Failures in nucleotide incorporation or block removal or incorporation of more than one nucleotide in a particular cycle&lt;br /&gt;
*Fading&lt;br /&gt;
**Decay in fluorescent signal intensity with each cycle&lt;br /&gt;
**Likely attributable to material loss during sequencing&lt;br /&gt;
*Crosstalk&lt;br /&gt;
**C channel illumination overlaps with A: a C label fluoresces in A channel (similarly G and T overlap)&lt;br /&gt;
**Likely caused by overlap in dye emission frequencies&lt;br /&gt;
*T Accumulation&lt;br /&gt;
**The fluorophores used for thymine are not always removed properly after each iteration&lt;br /&gt;
**Intensity of T signal increases across sequencing run&lt;br /&gt;
&lt;br /&gt;
==Alta-Cyclic==&lt;br /&gt;
===Training Stage===&lt;br /&gt;
*Learns run-specific noise patterns according to model and finds optimized solution reducing affect of noise sources using a Support Vector Machine (SVM)&lt;br /&gt;
*Half of training set used for cross-validation&lt;br /&gt;
&lt;br /&gt;
===Base Calling Stage===&lt;br /&gt;
*Reports all sequences from run with optimized parameters&lt;br /&gt;
&lt;br /&gt;
===Differences from Standard Illumina Base Caller===&lt;br /&gt;
*Calling parameters optimized empirically and tested to enhance accuracy of each run&lt;br /&gt;
*Calculates phasing parameters based on parametric model&lt;br /&gt;
*Dynamically tracks changes in crosstalk, which disrupt signals in later cycles&lt;br /&gt;
&lt;br /&gt;
==Probabilistic Base Calling==&lt;br /&gt;
*Produces an alternative probabilistic base calling method based on the fluorescence intensity quantifications that uses:&lt;br /&gt;
**Extended IUPAC alphabet to code ambiguous bases &lt;br /&gt;
**Information criterion to control length of trustable reads&lt;br /&gt;
*Reduced systematic bias by addressing:&lt;br /&gt;
**Crosstalk&lt;br /&gt;
**Dephasing&lt;br /&gt;
**Optical effect that tiles in center of image appear brighter corrected by fitting a 2D loess model to intensities and subtracting difference between fit and median intensities&lt;br /&gt;
*Measure level of uncertainty in base calling by entropy (uncertainty in determination of correct kth base)&lt;br /&gt;
*Does not consider fine-tuning image analysis&lt;br /&gt;
&lt;br /&gt;
==BayesCall==&lt;br /&gt;
*Model-based approach to base calling&lt;br /&gt;
*Main goal is to model sequencing process by taking stochasticity into account and by explicitly modeling how errors may arise&lt;br /&gt;
*Obtain base calls by maximizing posterior distribution of sequences given observed data and assuming a uniform prior on sequences&lt;br /&gt;
&lt;br /&gt;
==Swift==&lt;br /&gt;
Performs both image analysis and base calling&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Background subtraction – minimal pixel value within a window around each pixel subtracted from central pixel’s value&lt;br /&gt;
*Image correlation – alignment of images to reference cycle&lt;br /&gt;
*Object identification and intensity extraction&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Corrects for crosstalk by performing linear regression on crosstalk plots and use slope to derive correction matrix, performed iteratively until slope is zero&lt;br /&gt;
*Phasing correction by ranking clusters by chastity (the ratio of the highest intensity to the sum of the top two intensities) - use top 400 clusters to estimate phasing and apply it as a correction &lt;br /&gt;
*After correction, base with maximum intensity chosen as called base&lt;br /&gt;
&lt;br /&gt;
==Ibis==&lt;br /&gt;
&lt;br /&gt;
===Method===&lt;br /&gt;
*Estimate sequencing chemistry model as a parameter directly from data using statistical learning&lt;br /&gt;
*Training set from Bustard output using raw cluster intensities&lt;br /&gt;
*Used a base caller with SVM classifiers for each cycle that have intensity values of the current cycle as well as the previous and following cycles (if they exist)&lt;br /&gt;
*Data set created by aligning raw reads with mismatches for a fraction of the tiles to a reference sequence&lt;br /&gt;
**Half of this set used as a training set and the other half as a test set used to check results of training&lt;br /&gt;
*Estimate parameters for calculating a quality score given class assignment and distances to the classification/decision boundary from SVM&lt;br /&gt;
&lt;br /&gt;
===Comparison===&lt;br /&gt;
*Unlike AltaCyclic, includes base-specific phasing parameters so can correct raw intensities for T accumulation&lt;br /&gt;
*Does not call an &#039;N&#039; character for poor quality bases&lt;br /&gt;
*Process unique as causes of sequencing error not modeled separately&lt;br /&gt;
**Consider causes together by using neighboring signals in statistical learning procedure&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&lt;br /&gt;
Erlich, Y., Mitra, P.P., delaBastide, M., McCombie, W.R., Hannon, G.J. (2008) Alta-Cyclic: A self-optimizing base caller for next-generation sequencing. &#039;&#039;Nature Methods&#039;&#039; &#039;&#039;&#039;5&#039;&#039;&#039;:679-682 &lt;br /&gt;
&lt;br /&gt;
Kao, W.-C., Stevens, K., Song, Y.S. (2009) BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing. &#039;&#039;Genome Research&#039;&#039; &#039;&#039;&#039;19&#039;&#039;&#039;:1884-1895 &lt;br /&gt;
&lt;br /&gt;
Rougemont, J., Amzallag, A., Iseli, C., Farinelli, L., Xenarios, I., Naef, F. (2008) Probabilistic base calling of Solexa sequencing data. &#039;&#039;BMC Bioinformatics&#039;&#039; &#039;&#039;&#039;9&#039;&#039;&#039;:Article 431 &lt;br /&gt;
&lt;br /&gt;
Whiteford, N., Skelly, T., Curtis, C., Ritchie, M.E., Löhr, A., Zaranek, A.W., Abnizova, I., Brown, C. (2009) Swift: Primary data analysis for the Illumina Solexa sequencing platform. &#039;&#039;Bioinformatics&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:2194-2199&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=641</id>
		<title>Base Caller Summaries</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=641"/>
		<updated>2010-03-12T20:55:40Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: /* Ibis */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Standard Illumina Base Caller (Bustard)==&lt;br /&gt;
&lt;br /&gt;
===Sequencing-by-Synthesis (SBS)===&lt;br /&gt;
&lt;br /&gt;
*DNA sample obtained, containing many copies of same sequences and randomly fragmented&lt;br /&gt;
*Single-stranded DNA fragments attached to slide and amplified so there is a cluster of each fragment&lt;br /&gt;
*DNA polymerase and 4 terminal bases (with distinct fluorescent markers) added&lt;br /&gt;
*Clusters excited by lasers and photos taken in optimal wavelengths for 4 fluorophores&lt;br /&gt;
*Fluorophores and terminators removed and process repeated for L cycles&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Corrects for imperfect repositioning of camera and aberrations of lens by aligning images to reference from original cycle&lt;br /&gt;
*Signal for each cluster characterized as time series data of fluorescence intensities and noise&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Converts fluorescence signals into actual sequence data with quality scores&lt;br /&gt;
*Takes intensities of four channels for every cluster in each cycle and determines concentration of each base&lt;br /&gt;
*Renormalizes concentrations by multiplying by ratio of average concentrations in first cycle and current cycle&lt;br /&gt;
*Uses Markov model to determine transition matrix modeling probability of phasing (no new base synthesized), prephasing (two new bases synthesized), and normal incorporation&lt;br /&gt;
*Uses transition matrix and observed concentrations of each base to determine concentrations in absence of phasing and reports these as base calls&lt;br /&gt;
**Assumes crosstalk matrix constant for a given sequencing run and that phasing affects all nucleotides in the same way&lt;br /&gt;
&lt;br /&gt;
===General Noise Factors===&lt;br /&gt;
*Phasing&lt;br /&gt;
**Failures in nucleotide incorporation or block removal or incorporation of more than one nucleotide in a particular cycle&lt;br /&gt;
*Fading&lt;br /&gt;
**Decay in fluorescent signal intensity with each cycle&lt;br /&gt;
**Likely attributable to material loss during sequencing&lt;br /&gt;
*Crosstalk&lt;br /&gt;
**C channel illumination overlaps with A: a C label fluoresces in A channel (similarly G and T overlap)&lt;br /&gt;
**Likely caused by overlap in dye emission frequencies&lt;br /&gt;
*T Accumulation&lt;br /&gt;
**The fluorophores used for thymine are not always removed properly after each iteration&lt;br /&gt;
**Intensity of T signal increases across sequencing run&lt;br /&gt;
&lt;br /&gt;
==Alta-Cyclic==&lt;br /&gt;
===Training Stage===&lt;br /&gt;
*Learns run-specific noise patterns according to model and finds optimized solution reducing affect of noise sources using a Support Vector Machine (SVM)&lt;br /&gt;
*Half of training set used for cross-validation&lt;br /&gt;
&lt;br /&gt;
===Base Calling Stage===&lt;br /&gt;
*Reports all sequences from run with optimized parameters&lt;br /&gt;
&lt;br /&gt;
===Differences from Standard Illumina Base Caller===&lt;br /&gt;
*Calling parameters optimized empirically and tested to enhance accuracy of each run&lt;br /&gt;
*Calculates phasing parameters based on parametric model&lt;br /&gt;
*Dynamically tracks changes in crosstalk, which disrupt signals in later cycles&lt;br /&gt;
&lt;br /&gt;
==Probabilistic Base Calling==&lt;br /&gt;
*Produces an alternative probabilistic base calling method based on the fluorescence intensity quantifications that uses:&lt;br /&gt;
**Extended IUPAC alphabet to code ambiguous bases &lt;br /&gt;
**Information criterion to control length of trustable reads&lt;br /&gt;
*Reduced systematic bias by addressing:&lt;br /&gt;
**Crosstalk&lt;br /&gt;
**Dephasing&lt;br /&gt;
**Optical effect that tiles in center of image appear brighter corrected by fitting a 2D loess model to intensities and subtracting difference between fit and median intensities&lt;br /&gt;
*Measure level of uncertainty in base calling by entropy (uncertainty in determination of correct kth base)&lt;br /&gt;
*Does not consider fine-tuning image analysis&lt;br /&gt;
&lt;br /&gt;
==BayesCall==&lt;br /&gt;
*Model-based approach to base calling&lt;br /&gt;
*Main goal is to model sequencing process by taking stochasticity into account and by explicitly modeling how errors may arise&lt;br /&gt;
*Obtain base calls by maximizing posterior distribution of sequences given observed data and assuming a uniform prior on sequences&lt;br /&gt;
&lt;br /&gt;
==Swift==&lt;br /&gt;
Performs both image analysis and base calling&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Background subtraction – minimal pixel value within a window around each pixel subtracted from central pixel’s value&lt;br /&gt;
*Image correlation – alignment of images to reference cycle&lt;br /&gt;
*Object identification and intensity extraction&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Corrects for crosstalk by performing linear regression on crosstalk plots and use slope to derive correction matrix, performed iteratively until slope is zero&lt;br /&gt;
*Phasing correction by ranking clusters by chastity (the ratio of the highest intensity to the sum of the top two intensities) - use top 400 clusters to estimate phasing and apply it as a correction &lt;br /&gt;
*After correction, base with maximum intensity chosen as called base&lt;br /&gt;
&lt;br /&gt;
==Ibis==&lt;br /&gt;
&lt;br /&gt;
===Method===&lt;br /&gt;
*Estimate sequencing chemistry model as a parameter directly from data using statistical learning&lt;br /&gt;
*Training set from Bustard output using raw cluster intensities&lt;br /&gt;
*Used a base caller with SVM classifiers for each cycle that have intensity values of the current cycle as well as the previous and following cycles (if they exist)&lt;br /&gt;
*Data set created by aligning raw reads with mismatches for a fraction of the tiles to a reference sequence&lt;br /&gt;
**Half of this set used as a training set and the other half as a test set used to check results of training&lt;br /&gt;
*Estimate parameters for calculating a quality score given class assignment and distances to the classification/decision boundary from SVM&lt;br /&gt;
&lt;br /&gt;
===Comparison===&lt;br /&gt;
*Unlike AltaCyclic, includes base-specific phasing parameters so can correct raw intensities for T accumulation&lt;br /&gt;
*Does not call an &#039;N&#039; character for poor quality bases&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&lt;br /&gt;
Erlich, Y., Mitra, P.P., delaBastide, M., McCombie, W.R., Hannon, G.J. (2008) Alta-Cyclic: A self-optimizing base caller for next-generation sequencing. &#039;&#039;Nature Methods&#039;&#039; &#039;&#039;&#039;5&#039;&#039;&#039;:679-682 &lt;br /&gt;
&lt;br /&gt;
Kao, W.-C., Stevens, K., Song, Y.S. (2009) BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing. &#039;&#039;Genome Research&#039;&#039; &#039;&#039;&#039;19&#039;&#039;&#039;:1884-1895 &lt;br /&gt;
&lt;br /&gt;
Rougemont, J., Amzallag, A., Iseli, C., Farinelli, L., Xenarios, I., Naef, F. (2008) Probabilistic base calling of Solexa sequencing data. &#039;&#039;BMC Bioinformatics&#039;&#039; &#039;&#039;&#039;9&#039;&#039;&#039;:Article 431 &lt;br /&gt;
&lt;br /&gt;
Whiteford, N., Skelly, T., Curtis, C., Ritchie, M.E., Löhr, A., Zaranek, A.W., Abnizova, I., Brown, C. (2009) Swift: Primary data analysis for the Illumina Solexa sequencing platform. &#039;&#039;Bioinformatics&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:2194-2199&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=640</id>
		<title>Base Caller Summaries</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=640"/>
		<updated>2010-03-12T20:53:07Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: /* Ibis */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Standard Illumina Base Caller (Bustard)==&lt;br /&gt;
&lt;br /&gt;
===Sequencing-by-Synthesis (SBS)===&lt;br /&gt;
&lt;br /&gt;
*DNA sample obtained, containing many copies of same sequences and randomly fragmented&lt;br /&gt;
*Single-stranded DNA fragments attached to slide and amplified so there is a cluster of each fragment&lt;br /&gt;
*DNA polymerase and 4 terminal bases (with distinct fluorescent markers) added&lt;br /&gt;
*Clusters excited by lasers and photos taken in optimal wavelengths for 4 fluorophores&lt;br /&gt;
*Fluorophores and terminators removed and process repeated for L cycles&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Corrects for imperfect repositioning of camera and aberrations of lens by aligning images to reference from original cycle&lt;br /&gt;
*Signal for each cluster characterized as time series data of fluorescence intensities and noise&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Converts fluorescence signals into actual sequence data with quality scores&lt;br /&gt;
*Takes intensities of four channels for every cluster in each cycle and determines concentration of each base&lt;br /&gt;
*Renormalizes concentrations by multiplying by ratio of average concentrations in first cycle and current cycle&lt;br /&gt;
*Uses Markov model to determine transition matrix modeling probability of phasing (no new base synthesized), prephasing (two new bases synthesized), and normal incorporation&lt;br /&gt;
*Uses transition matrix and observed concentrations of each base to determine concentrations in absence of phasing and reports these as base calls&lt;br /&gt;
**Assumes crosstalk matrix constant for a given sequencing run and that phasing affects all nucleotides in the same way&lt;br /&gt;
&lt;br /&gt;
===General Noise Factors===&lt;br /&gt;
*Phasing&lt;br /&gt;
**Failures in nucleotide incorporation or block removal or incorporation of more than one nucleotide in a particular cycle&lt;br /&gt;
*Fading&lt;br /&gt;
**Decay in fluorescent signal intensity with each cycle&lt;br /&gt;
**Likely attributable to material loss during sequencing&lt;br /&gt;
*Crosstalk&lt;br /&gt;
**C channel illumination overlaps with A: a C label fluoresces in A channel (similarly G and T overlap)&lt;br /&gt;
**Likely caused by overlap in dye emission frequencies&lt;br /&gt;
*T Accumulation&lt;br /&gt;
**The fluorophores used for thymine are not always removed properly after each iteration&lt;br /&gt;
**Intensity of T signal increases across sequencing run&lt;br /&gt;
&lt;br /&gt;
==Alta-Cyclic==&lt;br /&gt;
===Training Stage===&lt;br /&gt;
*Learns run-specific noise patterns according to model and finds optimized solution reducing affect of noise sources using a Support Vector Machine (SVM)&lt;br /&gt;
*Half of training set used for cross-validation&lt;br /&gt;
&lt;br /&gt;
===Base Calling Stage===&lt;br /&gt;
*Reports all sequences from run with optimized parameters&lt;br /&gt;
&lt;br /&gt;
===Differences from Standard Illumina Base Caller===&lt;br /&gt;
*Calling parameters optimized empirically and tested to enhance accuracy of each run&lt;br /&gt;
*Calculates phasing parameters based on parametric model&lt;br /&gt;
*Dynamically tracks changes in crosstalk, which disrupt signals in later cycles&lt;br /&gt;
&lt;br /&gt;
==Probabilistic Base Calling==&lt;br /&gt;
*Produces an alternative probabilistic base calling method based on the fluorescence intensity quantifications that uses:&lt;br /&gt;
**Extended IUPAC alphabet to code ambiguous bases &lt;br /&gt;
**Information criterion to control length of trustable reads&lt;br /&gt;
*Reduced systematic bias by addressing:&lt;br /&gt;
**Crosstalk&lt;br /&gt;
**Dephasing&lt;br /&gt;
**Optical effect that tiles in center of image appear brighter corrected by fitting a 2D loess model to intensities and subtracting difference between fit and median intensities&lt;br /&gt;
*Measure level of uncertainty in base calling by entropy (uncertainty in determination of correct kth base)&lt;br /&gt;
*Does not consider fine-tuning image analysis&lt;br /&gt;
&lt;br /&gt;
==BayesCall==&lt;br /&gt;
*Model-based approach to base calling&lt;br /&gt;
*Main goal is to model sequencing process by taking stochasticity into account and by explicitly modeling how errors may arise&lt;br /&gt;
*Obtain base calls by maximizing posterior distribution of sequences given observed data and assuming a uniform prior on sequences&lt;br /&gt;
&lt;br /&gt;
==Swift==&lt;br /&gt;
Performs both image analysis and base calling&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Background subtraction – minimal pixel value within a window around each pixel subtracted from central pixel’s value&lt;br /&gt;
*Image correlation – alignment of images to reference cycle&lt;br /&gt;
*Object identification and intensity extraction&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Corrects for crosstalk by performing linear regression on crosstalk plots and use slope to derive correction matrix, performed iteratively until slope is zero&lt;br /&gt;
*Phasing correction by ranking clusters by chastity (the ratio of the highest intensity to the sum of the top two intensities) - use top 400 clusters to estimate phasing and apply it as a correction &lt;br /&gt;
*After correction, base with maximum intensity chosen as called base&lt;br /&gt;
&lt;br /&gt;
==Ibis==&lt;br /&gt;
&lt;br /&gt;
*Estimate sequencing chemistry model as a parameter directly from data using statistical learning&lt;br /&gt;
*Training set from Bustard output using raw cluster intensities&lt;br /&gt;
*Used a base caller with SVM classifiers for each cycle that have intensity values of the current cycle as well as the previous and following cycles (if they exist)&lt;br /&gt;
*Data set created by aligning raw reads with mismatches for a fraction of the tiles to a reference sequence&lt;br /&gt;
**Half of this set used as a training set and the other half as a test set used to check results of training&lt;br /&gt;
*Estimate parameters for calculating a quality score given class assignment and distances to the classification/decision boundary from SVM&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&lt;br /&gt;
Erlich, Y., Mitra, P.P., delaBastide, M., McCombie, W.R., Hannon, G.J. (2008) Alta-Cyclic: A self-optimizing base caller for next-generation sequencing. &#039;&#039;Nature Methods&#039;&#039; &#039;&#039;&#039;5&#039;&#039;&#039;:679-682 &lt;br /&gt;
&lt;br /&gt;
Kao, W.-C., Stevens, K., Song, Y.S. (2009) BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing. &#039;&#039;Genome Research&#039;&#039; &#039;&#039;&#039;19&#039;&#039;&#039;:1884-1895 &lt;br /&gt;
&lt;br /&gt;
Rougemont, J., Amzallag, A., Iseli, C., Farinelli, L., Xenarios, I., Naef, F. (2008) Probabilistic base calling of Solexa sequencing data. &#039;&#039;BMC Bioinformatics&#039;&#039; &#039;&#039;&#039;9&#039;&#039;&#039;:Article 431 &lt;br /&gt;
&lt;br /&gt;
Whiteford, N., Skelly, T., Curtis, C., Ritchie, M.E., Löhr, A., Zaranek, A.W., Abnizova, I., Brown, C. (2009) Swift: Primary data analysis for the Illumina Solexa sequencing platform. &#039;&#039;Bioinformatics&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:2194-2199&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=639</id>
		<title>Base Caller Summaries</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=639"/>
		<updated>2010-03-12T20:46:18Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: /* Standard Illumina Base Caller */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Standard Illumina Base Caller (Bustard)==&lt;br /&gt;
&lt;br /&gt;
===Sequencing-by-Synthesis (SBS)===&lt;br /&gt;
&lt;br /&gt;
*DNA sample obtained, containing many copies of same sequences and randomly fragmented&lt;br /&gt;
*Single-stranded DNA fragments attached to slide and amplified so there is a cluster of each fragment&lt;br /&gt;
*DNA polymerase and 4 terminal bases (with distinct fluorescent markers) added&lt;br /&gt;
*Clusters excited by lasers and photos taken in optimal wavelengths for 4 fluorophores&lt;br /&gt;
*Fluorophores and terminators removed and process repeated for L cycles&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Corrects for imperfect repositioning of camera and aberrations of lens by aligning images to reference from original cycle&lt;br /&gt;
*Signal for each cluster characterized as time series data of fluorescence intensities and noise&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Converts fluorescence signals into actual sequence data with quality scores&lt;br /&gt;
*Takes intensities of four channels for every cluster in each cycle and determines concentration of each base&lt;br /&gt;
*Renormalizes concentrations by multiplying by ratio of average concentrations in first cycle and current cycle&lt;br /&gt;
*Uses Markov model to determine transition matrix modeling probability of phasing (no new base synthesized), prephasing (two new bases synthesized), and normal incorporation&lt;br /&gt;
*Uses transition matrix and observed concentrations of each base to determine concentrations in absence of phasing and reports these as base calls&lt;br /&gt;
**Assumes crosstalk matrix constant for a given sequencing run and that phasing affects all nucleotides in the same way&lt;br /&gt;
&lt;br /&gt;
===General Noise Factors===&lt;br /&gt;
*Phasing&lt;br /&gt;
**Failures in nucleotide incorporation or block removal or incorporation of more than one nucleotide in a particular cycle&lt;br /&gt;
*Fading&lt;br /&gt;
**Decay in fluorescent signal intensity with each cycle&lt;br /&gt;
**Likely attributable to material loss during sequencing&lt;br /&gt;
*Crosstalk&lt;br /&gt;
**C channel illumination overlaps with A: a C label fluoresces in A channel (similarly G and T overlap)&lt;br /&gt;
**Likely caused by overlap in dye emission frequencies&lt;br /&gt;
*T Accumulation&lt;br /&gt;
**The fluorophores used for thymine are not always removed properly after each iteration&lt;br /&gt;
**Intensity of T signal increases across sequencing run&lt;br /&gt;
&lt;br /&gt;
==Alta-Cyclic==&lt;br /&gt;
===Training Stage===&lt;br /&gt;
*Learns run-specific noise patterns according to model and finds optimized solution reducing affect of noise sources using a Support Vector Machine (SVM)&lt;br /&gt;
*Half of training set used for cross-validation&lt;br /&gt;
&lt;br /&gt;
===Base Calling Stage===&lt;br /&gt;
*Reports all sequences from run with optimized parameters&lt;br /&gt;
&lt;br /&gt;
===Differences from Standard Illumina Base Caller===&lt;br /&gt;
*Calling parameters optimized empirically and tested to enhance accuracy of each run&lt;br /&gt;
*Calculates phasing parameters based on parametric model&lt;br /&gt;
*Dynamically tracks changes in crosstalk, which disrupt signals in later cycles&lt;br /&gt;
&lt;br /&gt;
==Probabilistic Base Calling==&lt;br /&gt;
*Produces an alternative probabilistic base calling method based on the fluorescence intensity quantifications that uses:&lt;br /&gt;
**Extended IUPAC alphabet to code ambiguous bases &lt;br /&gt;
**Information criterion to control length of trustable reads&lt;br /&gt;
*Reduced systematic bias by addressing:&lt;br /&gt;
**Crosstalk&lt;br /&gt;
**Dephasing&lt;br /&gt;
**Optical effect that tiles in center of image appear brighter corrected by fitting a 2D loess model to intensities and subtracting difference between fit and median intensities&lt;br /&gt;
*Measure level of uncertainty in base calling by entropy (uncertainty in determination of correct kth base)&lt;br /&gt;
*Does not consider fine-tuning image analysis&lt;br /&gt;
&lt;br /&gt;
==BayesCall==&lt;br /&gt;
*Model-based approach to base calling&lt;br /&gt;
*Main goal is to model sequencing process by taking stochasticity into account and by explicitly modeling how errors may arise&lt;br /&gt;
*Obtain base calls by maximizing posterior distribution of sequences given observed data and assuming a uniform prior on sequences&lt;br /&gt;
&lt;br /&gt;
==Swift==&lt;br /&gt;
Performs both image analysis and base calling&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Background subtraction – minimal pixel value within a window around each pixel subtracted from central pixel’s value&lt;br /&gt;
*Image correlation – alignment of images to reference cycle&lt;br /&gt;
*Object identification and intensity extraction&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Corrects for crosstalk by performing linear regression on crosstalk plots and use slope to derive correction matrix, performed iteratively until slope is zero&lt;br /&gt;
*Phasing correction by ranking clusters by chastity (the ratio of the highest intensity to the sum of the top two intensities) - use top 400 clusters to estimate phasing and apply it as a correction &lt;br /&gt;
*After correction, base with maximum intensity chosen as called base&lt;br /&gt;
&lt;br /&gt;
==Ibis==&lt;br /&gt;
&lt;br /&gt;
(To be added soon.)&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&lt;br /&gt;
Erlich, Y., Mitra, P.P., delaBastide, M., McCombie, W.R., Hannon, G.J. (2008) Alta-Cyclic: A self-optimizing base caller for next-generation sequencing. &#039;&#039;Nature Methods&#039;&#039; &#039;&#039;&#039;5&#039;&#039;&#039;:679-682 &lt;br /&gt;
&lt;br /&gt;
Kao, W.-C., Stevens, K., Song, Y.S. (2009) BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing. &#039;&#039;Genome Research&#039;&#039; &#039;&#039;&#039;19&#039;&#039;&#039;:1884-1895 &lt;br /&gt;
&lt;br /&gt;
Rougemont, J., Amzallag, A., Iseli, C., Farinelli, L., Xenarios, I., Naef, F. (2008) Probabilistic base calling of Solexa sequencing data. &#039;&#039;BMC Bioinformatics&#039;&#039; &#039;&#039;&#039;9&#039;&#039;&#039;:Article 431 &lt;br /&gt;
&lt;br /&gt;
Whiteford, N., Skelly, T., Curtis, C., Ritchie, M.E., Löhr, A., Zaranek, A.W., Abnizova, I., Brown, C. (2009) Swift: Primary data analysis for the Illumina Solexa sequencing platform. &#039;&#039;Bioinformatics&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:2194-2199&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=638</id>
		<title>Base Caller Summaries</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=638"/>
		<updated>2010-03-12T18:26:51Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: /* General Noise Factors */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Standard Illumina Base Caller==&lt;br /&gt;
&lt;br /&gt;
===Sequencing-by-Synthesis (SBS)===&lt;br /&gt;
&lt;br /&gt;
*DNA sample obtained, containing many copies of same sequences and randomly fragmented&lt;br /&gt;
*Single-stranded DNA fragments attached to slide and amplified so there is a cluster of each fragment&lt;br /&gt;
*DNA polymerase and 4 terminal bases (with distinct fluorescent markers) added&lt;br /&gt;
*Clusters excited by lasers and photos taken in optimal wavelengths for 4 fluorophores&lt;br /&gt;
*Fluorophores and terminators removed and process repeated for L cycles&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Corrects for imperfect repositioning of camera and aberrations of lens by aligning images to reference from original cycle&lt;br /&gt;
*Signal for each cluster characterized as time series data of fluorescence intensities and noise&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Converts fluorescence signals into actual sequence data with quality scores&lt;br /&gt;
*Takes intensities of four channels for every cluster in each cycle and determines concentration of each base&lt;br /&gt;
*Renormalizes concentrations by multiplying by ratio of average concentrations in first cycle and current cycle&lt;br /&gt;
*Uses Markov model to determine transition matrix modeling probability of phasing (no new base synthesized), prephasing (two new bases synthesized), and normal incorporation&lt;br /&gt;
*Uses transition matrix and observed concentrations of each base to determine concentrations in absence of phasing and reports these as base calls&lt;br /&gt;
&lt;br /&gt;
===General Noise Factors===&lt;br /&gt;
*Phasing&lt;br /&gt;
**Failures in nucleotide incorporation or block removal or incorporation of more than one nucleotide in a particular cycle&lt;br /&gt;
*Fading&lt;br /&gt;
**Decay in fluorescent signal intensity with each cycle&lt;br /&gt;
**Likely attributable to material loss during sequencing&lt;br /&gt;
*Crosstalk&lt;br /&gt;
**C channel illumination overlaps with A: a C label fluoresces in A channel (similarly G and T overlap)&lt;br /&gt;
**Likely caused by overlap in dye emission frequencies&lt;br /&gt;
*T Accumulation&lt;br /&gt;
**The fluorophores used for thymine are not always removed properly after each iteration&lt;br /&gt;
**Intensity of T signal increases across sequencing run&lt;br /&gt;
&lt;br /&gt;
==Alta-Cyclic==&lt;br /&gt;
===Training Stage===&lt;br /&gt;
*Learns run-specific noise patterns according to model and finds optimized solution reducing affect of noise sources using a Support Vector Machine (SVM)&lt;br /&gt;
*Half of training set used for cross-validation&lt;br /&gt;
&lt;br /&gt;
===Base Calling Stage===&lt;br /&gt;
*Reports all sequences from run with optimized parameters&lt;br /&gt;
&lt;br /&gt;
===Differences from Standard Illumina Base Caller===&lt;br /&gt;
*Calling parameters optimized empirically and tested to enhance accuracy of each run&lt;br /&gt;
*Calculates phasing parameters based on parametric model&lt;br /&gt;
*Dynamically tracks changes in crosstalk, which disrupt signals in later cycles&lt;br /&gt;
&lt;br /&gt;
==Probabilistic Base Calling==&lt;br /&gt;
*Produces an alternative probabilistic base calling method based on the fluorescence intensity quantifications that uses:&lt;br /&gt;
**Extended IUPAC alphabet to code ambiguous bases &lt;br /&gt;
**Information criterion to control length of trustable reads&lt;br /&gt;
*Reduced systematic bias by addressing:&lt;br /&gt;
**Crosstalk&lt;br /&gt;
**Dephasing&lt;br /&gt;
**Optical effect that tiles in center of image appear brighter corrected by fitting a 2D loess model to intensities and subtracting difference between fit and median intensities&lt;br /&gt;
*Measure level of uncertainty in base calling by entropy (uncertainty in determination of correct kth base)&lt;br /&gt;
*Does not consider fine-tuning image analysis&lt;br /&gt;
&lt;br /&gt;
==BayesCall==&lt;br /&gt;
*Model-based approach to base calling&lt;br /&gt;
*Main goal is to model sequencing process by taking stochasticity into account and by explicitly modeling how errors may arise&lt;br /&gt;
*Obtain base calls by maximizing posterior distribution of sequences given observed data and assuming a uniform prior on sequences&lt;br /&gt;
&lt;br /&gt;
==Swift==&lt;br /&gt;
Performs both image analysis and base calling&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Background subtraction – minimal pixel value within a window around each pixel subtracted from central pixel’s value&lt;br /&gt;
*Image correlation – alignment of images to reference cycle&lt;br /&gt;
*Object identification and intensity extraction&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Corrects for crosstalk by performing linear regression on crosstalk plots and use slope to derive correction matrix, performed iteratively until slope is zero&lt;br /&gt;
*Phasing correction by ranking clusters by chastity (the ratio of the highest intensity to the sum of the top two intensities) - use top 400 clusters to estimate phasing and apply it as a correction &lt;br /&gt;
*After correction, base with maximum intensity chosen as called base&lt;br /&gt;
&lt;br /&gt;
==Ibis==&lt;br /&gt;
&lt;br /&gt;
(To be added soon.)&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&lt;br /&gt;
Erlich, Y., Mitra, P.P., delaBastide, M., McCombie, W.R., Hannon, G.J. (2008) Alta-Cyclic: A self-optimizing base caller for next-generation sequencing. &#039;&#039;Nature Methods&#039;&#039; &#039;&#039;&#039;5&#039;&#039;&#039;:679-682 &lt;br /&gt;
&lt;br /&gt;
Kao, W.-C., Stevens, K., Song, Y.S. (2009) BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing. &#039;&#039;Genome Research&#039;&#039; &#039;&#039;&#039;19&#039;&#039;&#039;:1884-1895 &lt;br /&gt;
&lt;br /&gt;
Rougemont, J., Amzallag, A., Iseli, C., Farinelli, L., Xenarios, I., Naef, F. (2008) Probabilistic base calling of Solexa sequencing data. &#039;&#039;BMC Bioinformatics&#039;&#039; &#039;&#039;&#039;9&#039;&#039;&#039;:Article 431 &lt;br /&gt;
&lt;br /&gt;
Whiteford, N., Skelly, T., Curtis, C., Ritchie, M.E., Löhr, A., Zaranek, A.W., Abnizova, I., Brown, C. (2009) Swift: Primary data analysis for the Illumina Solexa sequencing platform. &#039;&#039;Bioinformatics&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:2194-2199&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=637</id>
		<title>Base Caller Summaries</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=637"/>
		<updated>2010-03-12T18:24:38Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: /* General Noise Factors */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Standard Illumina Base Caller==&lt;br /&gt;
&lt;br /&gt;
===Sequencing-by-Synthesis (SBS)===&lt;br /&gt;
&lt;br /&gt;
*DNA sample obtained, containing many copies of same sequences and randomly fragmented&lt;br /&gt;
*Single-stranded DNA fragments attached to slide and amplified so there is a cluster of each fragment&lt;br /&gt;
*DNA polymerase and 4 terminal bases (with distinct fluorescent markers) added&lt;br /&gt;
*Clusters excited by lasers and photos taken in optimal wavelengths for 4 fluorophores&lt;br /&gt;
*Fluorophores and terminators removed and process repeated for L cycles&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Corrects for imperfect repositioning of camera and aberrations of lens by aligning images to reference from original cycle&lt;br /&gt;
*Signal for each cluster characterized as time series data of fluorescence intensities and noise&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Converts fluorescence signals into actual sequence data with quality scores&lt;br /&gt;
*Takes intensities of four channels for every cluster in each cycle and determines concentration of each base&lt;br /&gt;
*Renormalizes concentrations by multiplying by ratio of average concentrations in first cycle and current cycle&lt;br /&gt;
*Uses Markov model to determine transition matrix modeling probability of phasing (no new base synthesized), prephasing (two new bases synthesized), and normal incorporation&lt;br /&gt;
*Uses transition matrix and observed concentrations of each base to determine concentrations in absence of phasing and reports these as base calls&lt;br /&gt;
&lt;br /&gt;
===General Noise Factors===&lt;br /&gt;
*Phasing&lt;br /&gt;
**Failures in nucleotide incorporation or block removal or incorporation of more than one nucleotide in a particular cycle&lt;br /&gt;
*Fading&lt;br /&gt;
**Decay in fluorescent signal intensity with each cycle&lt;br /&gt;
**Likely attributable to material loss during sequencing&lt;br /&gt;
*Crosstalk&lt;br /&gt;
**C channel illumination overlaps with A: a C label fluoresces in A channel (similarly G and T overlap)&lt;br /&gt;
**Likely caused by overlap in dye emission frequencies&lt;br /&gt;
*T Accumulation&lt;br /&gt;
**The fluorophores used for thymine show lower removal rate after each iteration and accumulate over the sequencing run&lt;br /&gt;
&lt;br /&gt;
==Alta-Cyclic==&lt;br /&gt;
===Training Stage===&lt;br /&gt;
*Learns run-specific noise patterns according to model and finds optimized solution reducing affect of noise sources using a Support Vector Machine (SVM)&lt;br /&gt;
*Half of training set used for cross-validation&lt;br /&gt;
&lt;br /&gt;
===Base Calling Stage===&lt;br /&gt;
*Reports all sequences from run with optimized parameters&lt;br /&gt;
&lt;br /&gt;
===Differences from Standard Illumina Base Caller===&lt;br /&gt;
*Calling parameters optimized empirically and tested to enhance accuracy of each run&lt;br /&gt;
*Calculates phasing parameters based on parametric model&lt;br /&gt;
*Dynamically tracks changes in crosstalk, which disrupt signals in later cycles&lt;br /&gt;
&lt;br /&gt;
==Probabilistic Base Calling==&lt;br /&gt;
*Produces an alternative probabilistic base calling method based on the fluorescence intensity quantifications that uses:&lt;br /&gt;
**Extended IUPAC alphabet to code ambiguous bases &lt;br /&gt;
**Information criterion to control length of trustable reads&lt;br /&gt;
*Reduced systematic bias by addressing:&lt;br /&gt;
**Crosstalk&lt;br /&gt;
**Dephasing&lt;br /&gt;
**Optical effect that tiles in center of image appear brighter corrected by fitting a 2D loess model to intensities and subtracting difference between fit and median intensities&lt;br /&gt;
*Measure level of uncertainty in base calling by entropy (uncertainty in determination of correct kth base)&lt;br /&gt;
*Does not consider fine-tuning image analysis&lt;br /&gt;
&lt;br /&gt;
==BayesCall==&lt;br /&gt;
*Model-based approach to base calling&lt;br /&gt;
*Main goal is to model sequencing process by taking stochasticity into account and by explicitly modeling how errors may arise&lt;br /&gt;
*Obtain base calls by maximizing posterior distribution of sequences given observed data and assuming a uniform prior on sequences&lt;br /&gt;
&lt;br /&gt;
==Swift==&lt;br /&gt;
Performs both image analysis and base calling&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Background subtraction – minimal pixel value within a window around each pixel subtracted from central pixel’s value&lt;br /&gt;
*Image correlation – alignment of images to reference cycle&lt;br /&gt;
*Object identification and intensity extraction&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Corrects for crosstalk by performing linear regression on crosstalk plots and use slope to derive correction matrix, performed iteratively until slope is zero&lt;br /&gt;
*Phasing correction by ranking clusters by chastity (the ratio of the highest intensity to the sum of the top two intensities) - use top 400 clusters to estimate phasing and apply it as a correction &lt;br /&gt;
*After correction, base with maximum intensity chosen as called base&lt;br /&gt;
&lt;br /&gt;
==Ibis==&lt;br /&gt;
&lt;br /&gt;
(To be added soon.)&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&lt;br /&gt;
Erlich, Y., Mitra, P.P., delaBastide, M., McCombie, W.R., Hannon, G.J. (2008) Alta-Cyclic: A self-optimizing base caller for next-generation sequencing. &#039;&#039;Nature Methods&#039;&#039; &#039;&#039;&#039;5&#039;&#039;&#039;:679-682 &lt;br /&gt;
&lt;br /&gt;
Kao, W.-C., Stevens, K., Song, Y.S. (2009) BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing. &#039;&#039;Genome Research&#039;&#039; &#039;&#039;&#039;19&#039;&#039;&#039;:1884-1895 &lt;br /&gt;
&lt;br /&gt;
Rougemont, J., Amzallag, A., Iseli, C., Farinelli, L., Xenarios, I., Naef, F. (2008) Probabilistic base calling of Solexa sequencing data. &#039;&#039;BMC Bioinformatics&#039;&#039; &#039;&#039;&#039;9&#039;&#039;&#039;:Article 431 &lt;br /&gt;
&lt;br /&gt;
Whiteford, N., Skelly, T., Curtis, C., Ritchie, M.E., Löhr, A., Zaranek, A.W., Abnizova, I., Brown, C. (2009) Swift: Primary data analysis for the Illumina Solexa sequencing platform. &#039;&#039;Bioinformatics&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:2194-2199&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=538</id>
		<title>Base Caller Summaries</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=538"/>
		<updated>2010-02-26T16:09:44Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Standard Illumina Base Caller==&lt;br /&gt;
&lt;br /&gt;
===Sequencing-by-Synthesis (SBS)===&lt;br /&gt;
&lt;br /&gt;
*DNA sample obtained, containing many copies of same sequences and randomly fragmented&lt;br /&gt;
*Single-stranded DNA fragments attached to slide and amplified so there is a cluster of each fragment&lt;br /&gt;
*DNA polymerase and 4 terminal bases (with distinct fluorescent markers) added&lt;br /&gt;
*Clusters excited by lasers and photos taken in optimal wavelengths for 4 fluorophores&lt;br /&gt;
*Fluorophores and terminators removed and process repeated for L cycles&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Corrects for imperfect repositioning of camera and aberrations of lens by aligning images to reference from original cycle&lt;br /&gt;
*Signal for each cluster characterized as time series data of fluorescence intensities and noise&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Converts fluorescence signals into actual sequence data with quality scores&lt;br /&gt;
*Takes intensities of four channels for every cluster in each cycle and determines concentration of each base&lt;br /&gt;
*Renormalizes concentrations by multiplying by ratio of average concentrations in first cycle and current cycle&lt;br /&gt;
*Uses Markov model to determine transition matrix modeling probability of phasing (no new base synthesized), prephasing (two new bases synthesized), and normal incorporation&lt;br /&gt;
*Uses transition matrix and observed concentrations of each base to determine concentrations in absence of phasing and reports these as base calls&lt;br /&gt;
&lt;br /&gt;
===General Noise Factors===&lt;br /&gt;
*Phasing&lt;br /&gt;
**Failures in nucleotide incorporation or block removal or incorporation of more than one nucleotide in a particular cycle&lt;br /&gt;
*Fading&lt;br /&gt;
**Decay in fluorescent signal intensity with each cycle&lt;br /&gt;
**Likely attributable to material loss during sequencing&lt;br /&gt;
*Crosstalk&lt;br /&gt;
**C channel illumination overlaps with A: a C label fluoresces in A channel (similarly G and T overlap)&lt;br /&gt;
**Likely caused by overlap in dye emission frequencies &lt;br /&gt;
&lt;br /&gt;
==Alta-Cyclic==&lt;br /&gt;
===Training Stage===&lt;br /&gt;
*Learns run-specific noise patterns according to model and finds optimized solution reducing affect of noise sources using a Support Vector Machine (SVM)&lt;br /&gt;
*Half of training set used for cross-validation&lt;br /&gt;
&lt;br /&gt;
===Base Calling Stage===&lt;br /&gt;
*Reports all sequences from run with optimized parameters&lt;br /&gt;
&lt;br /&gt;
===Differences from Standard Illumina Base Caller===&lt;br /&gt;
*Calling parameters optimized empirically and tested to enhance accuracy of each run&lt;br /&gt;
*Calculates phasing parameters based on parametric model&lt;br /&gt;
*Dynamically tracks changes in crosstalk, which disrupt signals in later cycles&lt;br /&gt;
&lt;br /&gt;
==Probabilistic Base Calling==&lt;br /&gt;
*Produces an alternative probabilistic base calling method based on the fluorescence intensity quantifications that uses:&lt;br /&gt;
**Extended IUPAC alphabet to code ambiguous bases &lt;br /&gt;
**Information criterion to control length of trustable reads&lt;br /&gt;
*Reduced systematic bias by addressing:&lt;br /&gt;
**Crosstalk&lt;br /&gt;
**Dephasing&lt;br /&gt;
**Optical effect that tiles in center of image appear brighter corrected by fitting a 2D loess model to intensities and subtracting difference between fit and median intensities&lt;br /&gt;
*Measure level of uncertainty in base calling by entropy (uncertainty in determination of correct kth base)&lt;br /&gt;
*Does not consider fine-tuning image analysis&lt;br /&gt;
&lt;br /&gt;
==BayesCall==&lt;br /&gt;
*Model-based approach to base calling&lt;br /&gt;
*Main goal is to model sequencing process by taking stochasticity into account and by explicitly modeling how errors may arise&lt;br /&gt;
*Obtain base calls by maximizing posterior distribution of sequences given observed data and assuming a uniform prior on sequences&lt;br /&gt;
&lt;br /&gt;
==Swift==&lt;br /&gt;
Performs both image analysis and base calling&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Background subtraction – minimal pixel value within a window around each pixel subtracted from central pixel’s value&lt;br /&gt;
*Image correlation – alignment of images to reference cycle&lt;br /&gt;
*Object identification and intensity extraction&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Corrects for crosstalk by performing linear regression on crosstalk plots and use slope to derive correction matrix, performed iteratively until slope is zero&lt;br /&gt;
*Phasing correction by ranking clusters by chastity (the ratio of the highest intensity to the sum of the top two intensities) - use top 400 clusters to estimate phasing and apply it as a correction &lt;br /&gt;
*After correction, base with maximum intensity chosen as called base&lt;br /&gt;
&lt;br /&gt;
==Ibis==&lt;br /&gt;
&lt;br /&gt;
(To be added soon.)&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&lt;br /&gt;
Erlich, Y., Mitra, P.P., delaBastide, M., McCombie, W.R., Hannon, G.J. (2008) Alta-Cyclic: A self-optimizing base caller for next-generation sequencing. &#039;&#039;Nature Methods&#039;&#039; &#039;&#039;&#039;5&#039;&#039;&#039;:679-682 &lt;br /&gt;
&lt;br /&gt;
Kao, W.-C., Stevens, K., Song, Y.S. (2009) BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing. &#039;&#039;Genome Research&#039;&#039; &#039;&#039;&#039;19&#039;&#039;&#039;:1884-1895 &lt;br /&gt;
&lt;br /&gt;
Rougemont, J., Amzallag, A., Iseli, C., Farinelli, L., Xenarios, I., Naef, F. (2008) Probabilistic base calling of Solexa sequencing data. &#039;&#039;BMC Bioinformatics&#039;&#039; &#039;&#039;&#039;9&#039;&#039;&#039;:Article 431 &lt;br /&gt;
&lt;br /&gt;
Whiteford, N., Skelly, T., Curtis, C., Ritchie, M.E., Löhr, A., Zaranek, A.W., Abnizova, I., Brown, C. (2009) Swift: Primary data analysis for the Illumina Solexa sequencing platform. &#039;&#039;Bioinformatics&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:2194-2199&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=537</id>
		<title>Base Caller Summaries</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=537"/>
		<updated>2010-02-26T16:04:50Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Standard Illumina Base Caller==&lt;br /&gt;
&lt;br /&gt;
===Sequencing-by-Synthesis (SBS)===&lt;br /&gt;
&lt;br /&gt;
*DNA sample obtained, containing many copies of same sequences and randomly fragmented&lt;br /&gt;
*Single-stranded DNA fragments attached to slide and amplified so there is a cluster of each fragment&lt;br /&gt;
*DNA polymerase and 4 terminal bases (with distinct fluorescent markers) added&lt;br /&gt;
*Clusters excited by lasers and photos taken in optimal wavelengths for 4 fluorophores&lt;br /&gt;
*Fluorophores and terminators removed and process repeated for L cycles&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Corrects for imperfect repositioning of camera and aberrations of lens by aligning images to reference from original cycle&lt;br /&gt;
*Signal for each cluster characterized as time series data of fluorescence intensities and noise&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Converts fluorescence signals into actual sequence data with quality scores&lt;br /&gt;
*Takes intensities of four channels for every cluster in each cycle and determines concentration of each base&lt;br /&gt;
*Renormalizes concentrations by multiplying by ratio of average concentrations in first cycle and current cycle&lt;br /&gt;
*Uses Markov model to determine transition matrix modeling probability of phasing (no new base synthesized), prephasing (two new bases synthesized), and normal incorporation&lt;br /&gt;
*Uses transition matrix and observed concentrations of each base to determine concentrations in absence of phasing and reports these as base calls&lt;br /&gt;
&lt;br /&gt;
===General Noise Factors===&lt;br /&gt;
*Phasing&lt;br /&gt;
**Failures in nucleotide incorporation or block removal or incorporation of more than one nucleotide in a particular cycle&lt;br /&gt;
*Fading&lt;br /&gt;
**Decay in fluorescent signal intensity with each cycle&lt;br /&gt;
**Likely attributable to material loss during sequencing&lt;br /&gt;
*Crosstalk&lt;br /&gt;
**C channel illumination overlaps with A: a C label fluoresces in A channel (similarly G and T overlap)&lt;br /&gt;
**Likely caused by overlap in dye emission frequencies &lt;br /&gt;
&lt;br /&gt;
==Alta-Cyclic==&lt;br /&gt;
===Training Stage===&lt;br /&gt;
*Learns run-specific noise patterns according to model and finds optimized solution reducing affect of noise sources using a Support Vector Machine (SVM)&lt;br /&gt;
*Half of training set used for cross-validation&lt;br /&gt;
&lt;br /&gt;
===Base Calling Stage===&lt;br /&gt;
*Reports all sequences from run with optimized parameters&lt;br /&gt;
&lt;br /&gt;
===Differences from Standard Illumina Base Caller===&lt;br /&gt;
*Calling parameters optimized empirically and tested to enhance accuracy of each run&lt;br /&gt;
*Calculates phasing parameters based on parametric model&lt;br /&gt;
*Dynamically tracks changes in crosstalk, which disrupt signals in later cycles&lt;br /&gt;
&lt;br /&gt;
==Probabilistic Base Calling==&lt;br /&gt;
&lt;br /&gt;
==BayesCall==&lt;br /&gt;
&lt;br /&gt;
==Swift==&lt;br /&gt;
&lt;br /&gt;
==Ibis==&lt;br /&gt;
&lt;br /&gt;
(To be added soon.)&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&lt;br /&gt;
Erlich, Y., Mitra, P.P., delaBastide, M., McCombie, W.R., Hannon, G.J. (2008) Alta-Cyclic: A self-optimizing base caller for next-generation sequencing. &#039;&#039;Nature Methods&#039;&#039; &#039;&#039;&#039;5&#039;&#039;&#039;:679-682 &lt;br /&gt;
&lt;br /&gt;
Kao, W.-C., Stevens, K., Song, Y.S. (2009) BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing. &#039;&#039;Genome Research&#039;&#039; &#039;&#039;&#039;19&#039;&#039;&#039;:1884-1895 &lt;br /&gt;
&lt;br /&gt;
Rougemont, J., Amzallag, A., Iseli, C., Farinelli, L., Xenarios, I., Naef, F. (2008) Probabilistic base calling of Solexa sequencing data. &#039;&#039;BMC Bioinformatics&#039;&#039; &#039;&#039;&#039;9&#039;&#039;&#039;:Article 431 &lt;br /&gt;
&lt;br /&gt;
Whiteford, N., Skelly, T., Curtis, C., Ritchie, M.E., Löhr, A., Zaranek, A.W., Abnizova, I., Brown, C. (2009) Swift: Primary data analysis for the Illumina Solexa sequencing platform. &#039;&#039;Bioinformatics&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:2194-2199&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=536</id>
		<title>Base Caller Summaries</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=536"/>
		<updated>2010-02-26T16:01:54Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Standard Illumina Base Caller==&lt;br /&gt;
&lt;br /&gt;
===Sequencing-by-Synthesis (SBS)===&lt;br /&gt;
&lt;br /&gt;
*DNA sample obtained, containing many copies of same sequences and randomly fragmented&lt;br /&gt;
*Single-stranded DNA fragments attached to slide and amplified so there is a cluster of each fragment&lt;br /&gt;
*DNA polymerase and 4 terminal bases (with distinct fluorescent markers) added&lt;br /&gt;
*Clusters excited by lasers and photos taken in optimal wavelengths for 4 fluorophores&lt;br /&gt;
*Fluorophores and terminators removed and process repeated for L cycles&lt;br /&gt;
&lt;br /&gt;
===Image Analysis===&lt;br /&gt;
*Corrects for imperfect repositioning of camera and aberrations of lens by aligning images to reference from original cycle&lt;br /&gt;
*Signal for each cluster characterized as time series data of fluorescence intensities and noise&lt;br /&gt;
&lt;br /&gt;
===Base Calling===&lt;br /&gt;
*Converts fluorescence signals into actual sequence data with quality scores&lt;br /&gt;
*Takes intensities of four channels for every cluster in each cycle and determines concentration of each base&lt;br /&gt;
*Renormalizes concentrations by multiplying by ratio of average concentrations in first cycle and current cycle&lt;br /&gt;
*Uses Markov model to determine transition matrix modeling probability of phasing (no new base synthesized), prephasing (two new bases synthesized), and normal incorporation&lt;br /&gt;
*Uses transition matrix and observed concentrations of each base to determine concentrations in absence of phasing and reports these as base calls&lt;br /&gt;
&lt;br /&gt;
===General Noise Factors===&lt;br /&gt;
*Phasing&lt;br /&gt;
**Failures in nucleotide incorporation or block removal or incorporation of more than one nucleotide in a particular cycle&lt;br /&gt;
*Fading&lt;br /&gt;
**Decay in fluorescent signal intensity with each cycle&lt;br /&gt;
**Likely attributable to material loss during sequencing&lt;br /&gt;
*Crosstalk&lt;br /&gt;
**C channel illumination overlaps with A: a C label fluoresces in A channel (similarly G and T overlap)&lt;br /&gt;
**Likely caused by overlap in dye emission frequencies &lt;br /&gt;
&lt;br /&gt;
==Alta-Cyclic==&lt;br /&gt;
&lt;br /&gt;
==Probabilistic Base Calling==&lt;br /&gt;
&lt;br /&gt;
==BayesCall==&lt;br /&gt;
&lt;br /&gt;
==Swift==&lt;br /&gt;
&lt;br /&gt;
==Ibis==&lt;br /&gt;
&lt;br /&gt;
(To be added soon.)&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&lt;br /&gt;
Erlich, Y., Mitra, P.P., delaBastide, M., McCombie, W.R., Hannon, G.J. (2008) Alta-Cyclic: A self-optimizing base caller for next-generation sequencing. &#039;&#039;Nature Methods&#039;&#039; &#039;&#039;&#039;5&#039;&#039;&#039;:679-682 &lt;br /&gt;
&lt;br /&gt;
Kao, W.-C., Stevens, K., Song, Y.S. (2009) BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing. &#039;&#039;Genome Research&#039;&#039; &#039;&#039;&#039;19&#039;&#039;&#039;:1884-1895 &lt;br /&gt;
&lt;br /&gt;
Rougemont, J., Amzallag, A., Iseli, C., Farinelli, L., Xenarios, I., Naef, F. (2008) Probabilistic base calling of Solexa sequencing data. &#039;&#039;BMC Bioinformatics&#039;&#039; &#039;&#039;&#039;9&#039;&#039;&#039;:Article 431 &lt;br /&gt;
&lt;br /&gt;
Whiteford, N., Skelly, T., Curtis, C., Ritchie, M.E., Löhr, A., Zaranek, A.W., Abnizova, I., Brown, C. (2009) Swift: Primary data analysis for the Illumina Solexa sequencing platform. &#039;&#039;Bioinformatics&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:2194-2199&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=535</id>
		<title>Base Caller Summaries</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=535"/>
		<updated>2010-02-26T15:49:38Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Standard Illumina Base Caller =&lt;br /&gt;
&lt;br /&gt;
= Alta-Cyclic =&lt;br /&gt;
&lt;br /&gt;
= Probabilistic Base Calling =&lt;br /&gt;
&lt;br /&gt;
= BayesCall =&lt;br /&gt;
&lt;br /&gt;
= Swift =&lt;br /&gt;
&lt;br /&gt;
= Ibis =&lt;br /&gt;
&lt;br /&gt;
(To be added soon.)&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
&lt;br /&gt;
Erlich, Y., Mitra, P.P., delaBastide, M., McCombie, W.R., Hannon, G.J. (2008) Alta-Cyclic: A self-optimizing base caller for next-generation sequencing. &#039;&#039;Nature Methods&#039;&#039; &#039;&#039;&#039;5&#039;&#039;&#039;:679-682 &lt;br /&gt;
&lt;br /&gt;
Kao, W.-C., Stevens, K., Song, Y.S. (2009) BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing. &#039;&#039;Genome Research&#039;&#039; &#039;&#039;&#039;19&#039;&#039;&#039;:1884-1895 &lt;br /&gt;
&lt;br /&gt;
Rougemont, J., Amzallag, A., Iseli, C., Farinelli, L., Xenarios, I., Naef, F. (2008) Probabilistic base calling of Solexa sequencing data. &#039;&#039;BMC Bioinformatics&#039;&#039; &#039;&#039;&#039;9&#039;&#039;&#039;:Article 431 &lt;br /&gt;
&lt;br /&gt;
Whiteford, N., Skelly, T., Curtis, C., Ritchie, M.E., Löhr, A., Zaranek, A.W., Abnizova, I., Brown, C. (2009) Swift: Primary data analysis for the Illumina Solexa sequencing platform. &#039;&#039;Bioinformatics&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:2194-2199&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=534</id>
		<title>Base Caller Summaries</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Base_Caller_Summaries&amp;diff=534"/>
		<updated>2010-02-26T15:45:19Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: Created page with &amp;#039;= Standard Illumina Base Caller =  = Alta-Cyclic =  = Probabilistic Base Calling =  = BayesCall =  = Swift =  = IBIS =  (To be added soon.)  = References =&amp;#039;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Standard Illumina Base Caller =&lt;br /&gt;
&lt;br /&gt;
= Alta-Cyclic =&lt;br /&gt;
&lt;br /&gt;
= Probabilistic Base Calling =&lt;br /&gt;
&lt;br /&gt;
= BayesCall =&lt;br /&gt;
&lt;br /&gt;
= Swift =&lt;br /&gt;
&lt;br /&gt;
= IBIS =&lt;br /&gt;
&lt;br /&gt;
(To be added soon.)&lt;br /&gt;
&lt;br /&gt;
= References =&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Links_to_Sequence_Analysis_Tools&amp;diff=533</id>
		<title>Links to Sequence Analysis Tools</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Links_to_Sequence_Analysis_Tools&amp;diff=533"/>
		<updated>2010-02-26T15:42:29Z</updated>

		<summary type="html">&lt;p&gt;Srashkin: /* Base Calling Tools */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[http://www.broadinstitute.org/gsa/wiki/index.php/Base_quality_score_recalibration Quality Score Calibration Tool]] at the Broad Institute.&lt;br /&gt;
&lt;br /&gt;
== Base Calling Tools ==&lt;br /&gt;
&lt;br /&gt;
[[Base Caller Summaries]]&lt;br /&gt;
&lt;br /&gt;
Erlich, Y., Mitra, P.P., delaBastide, M., McCombie, W.R., Hannon, G.J. (2008) Alta-Cyclic: A self-optimizing base caller for next-generation sequencing. &#039;&#039;Nature Methods&#039;&#039; &#039;&#039;&#039;5&#039;&#039;&#039;:679-682&lt;br /&gt;
&lt;br /&gt;
Rougemont, J., Amzallag, A., Iseli, C., Farinelli, L., Xenarios, I., Naef, F. (2008) Probabilistic base calling of Solexa sequencing data. &#039;&#039;BMC Bioinformatics&#039;&#039; &#039;&#039;&#039;9&#039;&#039;&#039;:Article 431&lt;br /&gt;
&lt;br /&gt;
Kao, W.-C., Stevens, K., Song, Y.S. (2009) BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing. &#039;&#039;Genome Research&#039;&#039; &#039;&#039;&#039;19&#039;&#039;&#039;:1884-1895&lt;br /&gt;
&lt;br /&gt;
Whiteford, N., Skelly, T., Curtis, C., Ritchie, M.E., Löhr, A., Zaranek, A.W., Abnizova, I., Brown, C. (2009) &lt;br /&gt;
Swift: Primary data analysis for the Illumina Solexa sequencing platform. &#039;&#039;Bioinformatics&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:2194-2199&lt;/div&gt;</summary>
		<author><name>Srashkin</name></author>
	</entry>
</feed>