<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://wiki.ubc.ca/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Ezhao</id>
	<title>UBC Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.ubc.ca/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Ezhao"/>
	<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/Special:Contributions/Ezhao"/>
	<updated>2026-08-01T16:34:53Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.43.9</generator>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Thread:User_talk:Ezhao/Moved_your_new_page_to_the_Sandbox/reply&amp;diff=95138</id>
		<title>Thread:User talk:Ezhao/Moved your new page to the Sandbox/reply</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Thread:User_talk:Ezhao/Moved_your_new_page_to_the_Sandbox/reply&amp;diff=95138"/>
		<updated>2011-05-18T21:26:30Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: Reply to Moved your new page to the Sandbox&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Ah, perfect.  I was wondering where the place was for that kind of thing.  I thought I would be able to delete the page.  Thanks for your help,&lt;br /&gt;
&lt;br /&gt;
Eric&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Sandbox:Best_Practice_Variant_Detection_with_the_GATK_v2&amp;diff=95015</id>
		<title>Sandbox:Best Practice Variant Detection with the GATK v2</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Sandbox:Best_Practice_Variant_Detection_with_the_GATK_v2&amp;diff=95015"/>
		<updated>2011-05-18T16:45:41Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: Blanked the page&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Titv_expectations.jpg&amp;diff=95013</id>
		<title>File:Titv expectations.jpg</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Titv_expectations.jpg&amp;diff=95013"/>
		<updated>2011-05-18T16:44:45Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Best_practice_calls.jpg&amp;diff=95012</id>
		<title>File:Best practice calls.jpg</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Best_practice_calls.jpg&amp;diff=95012"/>
		<updated>2011-05-18T16:44:34Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:DataProcessingPPL.jpg&amp;diff=95011</id>
		<title>File:DataProcessingPPL.jpg</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:DataProcessingPPL.jpg&amp;diff=95011"/>
		<updated>2011-05-18T16:43:57Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Overall_flow.jpg&amp;diff=95010</id>
		<title>File:Overall flow.jpg</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Overall_flow.jpg&amp;diff=95010"/>
		<updated>2011-05-18T16:43:40Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Sandbox:Best_Practice_Variant_Detection_with_the_GATK_v2&amp;diff=95009</id>
		<title>Sandbox:Best Practice Variant Detection with the GATK v2</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Sandbox:Best_Practice_Variant_Detection_with_the_GATK_v2&amp;diff=95009"/>
		<updated>2011-05-18T16:41:50Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Data Processing Pipeline Script ==&lt;br /&gt;
The [[Data Processing Pipeline]] is a Queue script that performs all data processing described in this page following our most current recommendations for best practices.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
Our current best practice for making SNP calls is divided into 5 sequential steps (see Figure): initial mapping, refinement of the initial reads, multi-sample indel and SNP calling, filtering of the raw SNP calls, and finally variant quality score recalibration.  These steps are the same for targeted resequencing, whole exomes, deep whole genomes, and low-pass whole genomes.  The exact commands for each tool are available on the tools wiki entry.   &lt;br /&gt;
&lt;br /&gt;
Note that, due to the specific attributes of a project, that the the specific values used in each of the commands may need to be selected by the analyst.  Care should be taken by the analyst running our tools to understand what each parameter does and to evaluate which value best fits his/her data.&lt;br /&gt;
&lt;br /&gt;
[[File:overall_flow.jpg|300px|thumb|Overview of GATK data processing pipeline]]&lt;br /&gt;
&lt;br /&gt;
=== Lanes, Samples, Cohort ===&lt;br /&gt;
&lt;br /&gt;
There are four major organizational units for next-generation DNA sequencing processes:&lt;br /&gt;
&lt;br /&gt;
; Lane : The basic machine unit for sequencing.  The lane reflects the basic independent run of an NGS machine.  For Illumina machines, this is the physical sequencing lane.  &lt;br /&gt;
&lt;br /&gt;
; Library : A unit of DNA preparation that at some point is physically pooled together.  Multiple lanes can be run from aliquots from the same library.  The DNA library and its preparation is the natural unit that is being sequenced.  For example, if the library has limited complexity, then many sequences are duplicated and will result in a high duplication rate across lanes. &lt;br /&gt;
&lt;br /&gt;
; Sample : A single individual, such as human CEPH NA12878.  Multiple libraries with different properties can be constructed from the original sample DNA source.  Here we treat samples as independent individuals whose genome sequence we are attempting to determine.  From this perspective, tumor / normal samples are different despite coming from the same individual.&lt;br /&gt;
&lt;br /&gt;
; Cohort : A collection of samples being analyzed together.  This organizational unit is the most subjective and depends intimately on the design goals of the sequencing project.  For population discovery projects like the 1000 Genomes, the analysis cohort is the ~100 individual in each population.  For exome projects with many samples (e.g., ESP with 800 EOMI samples) deeply sequenced we divide up the complete set of samples into cohorts of ~50 individuals for multi-sample analyses.  &lt;br /&gt;
&lt;br /&gt;
This document describes how to call variation within a single analysis cohort, comprised for one or many samples, each of one or many libraries that were sequenced on at least one lane of an NGS machine. &lt;br /&gt;
&lt;br /&gt;
Note that many GATK commands can be run at the lane level, but will give better results seeing all of the data for a single sample, or even all of the data for all samples.  Unfortunately, there&#039;s a trade-off in computational cost by running these commands across all of your data simultaneously.&lt;br /&gt;
&lt;br /&gt;
=== Testing data: 64x HiSeq on chr20 for NA12878 ===&lt;br /&gt;
&lt;br /&gt;
In order to help individuals get up to speed, evaluate their command lines, and generally become familiar with the GATK tools we recommend you download the raw and realigned, recalibrated [[NA12878 test data]].  It should be possible to apply all of the approaches outlined below to get excellent results for realignment, recalibration, SNP calling, indel calling, filtering, and variant quality score recalibration using this data.&lt;br /&gt;
&lt;br /&gt;
== Phase I: Raw data processing ==&lt;br /&gt;
[[Image:DataProcessingPPL.jpg|thumb|300px|Data processing pipeline of the best practices for raw data processing, from sequencer data (fastq files) to analysis read reads (bam file)]]&lt;br /&gt;
&lt;br /&gt;
=== Initial read mapping ===&lt;br /&gt;
&lt;br /&gt;
The GATK data processing pipeline assumes that one of the many NGS read aligners (see [http://bib.oxfordjournals.org/content/11/5/473.abstract] for a review) has been applied to your raw FASTQ files.  For Illumina data we recommend [http://bio-bwa.sourceforge.net/ BWA] because it is accurate, fast, well-supported, open-source, and emits BAM files natively.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Raw BAM to realigned, recalibrated BAM ===&lt;br /&gt;
&lt;br /&gt;
The three key tools here are [[Base quality score recalibration]], [[Local realignment around indels]], and [http://picard.sourceforge.net/command-line-overview.shtml#MarkDuplicates|Picard&#039;s MarkDuplicates].  Although ideally one follows the flow from Figure 1, in practice MarkDuplicates can be run before local realignment, in order to handle cases where duplicates overlapping indels get marginally different alignments   (unlikely but possible) and so will not be considered as potential duplicates because MarkDuplicates looks only at read pair start and stop positions.  There are several options here, from the easy and fast basic protocol to the more.&lt;br /&gt;
&lt;br /&gt;
There are two types of realignment:&lt;br /&gt;
* Realignment only at known sites, which is very efficient, can operate with little coverage (1x per lane genome wide) but can only realign reads at known indels.&lt;br /&gt;
* Fully local realignment uses mismatching bases to determine if a site should be realigned, and relies on sufficient coverage to discover the correct indel allele in the reads for alignment.  It is much slower (involves SW step) but can discover new indel sites in the reads.  If you have a database of known indels (for human, this database is extensive) then at this stage you would also include these indels during realignment, which vastly improves sensitivity, specificity, and speed.&lt;br /&gt;
&lt;br /&gt;
==== Previous recommendation: lane-level recalibration, sample-level realignment ====&lt;br /&gt;
&lt;br /&gt;
This is the protocol we used at the Broad Institute for the last year.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each lane.bam&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicate(lane.bam)&lt;br /&gt;
    recal.bam &amp;lt;- recal(dedup.bam)&lt;br /&gt;
&lt;br /&gt;
for each sample&lt;br /&gt;
    lanes.bam &amp;lt;- merged recal.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(lanes.bam)&lt;br /&gt;
    realigned.bam &amp;lt;- realign(dedup.bam) [with known sites if possible]&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fast: lane-level realignment at known sites only and lane-level recalibration ====&lt;br /&gt;
&lt;br /&gt;
This protocol adds lane-level local realignment around known indels, making it very fast (no sample level processing) and gives better results for human samples than the previous recommendation:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each lane.bam&lt;br /&gt;
    realigned.bam &amp;lt;- realign(lane.bam) [at only known sites]&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicate(realigned.bam)&lt;br /&gt;
    recal.bam &amp;lt;- recal(dedup.bam)&lt;br /&gt;
&lt;br /&gt;
for each sample&lt;br /&gt;
    recals.bam &amp;lt;- merged lane-level recal.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(recals.bam)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fast + sample-level realignment ====&lt;br /&gt;
&lt;br /&gt;
This protocol adds sample-level realignment after library / sample level dedupping, so that novel indels in each sample can be discovered and realigned around.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each lane.bam&lt;br /&gt;
    realigned.bam &amp;lt;- realign(lane.bam) [at only known sites]&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicate(realigned.bam)&lt;br /&gt;
    recal.bam &amp;lt;- recal(dedup.bam)&lt;br /&gt;
&lt;br /&gt;
for each sample&lt;br /&gt;
    recals.bam &amp;lt;- merged lane-level recal.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(recals.bam)&lt;br /&gt;
    realigned.bam &amp;lt;- realign(dedup.bam) [with known sites if possible]&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Better: sample-level realignment with known indels and recalibration ====&lt;br /&gt;
&lt;br /&gt;
Rather than doing the lane level cleaning and recalibration, this process aggregates all of the reads for each sample and then does a full dedupping, realign, and recalibration, yielding the best single-sample results.  The big change here is sample-level cleaning followed by recalibration, giving you the most accurate quality scores possible for a single sample. &lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each sample&lt;br /&gt;
    lanes.bam &amp;lt;- merged lane.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(lanes.bam)&lt;br /&gt;
    realigned.bam &amp;lt;- realign(dedup.bam) [with known sites if possible]&lt;br /&gt;
    recal.bam &amp;lt;- recal(realigned.bam)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Best: multi-sample realignment with known sites and recalibration ====&lt;br /&gt;
&lt;br /&gt;
Finally, if you really want to get the absolute best results, whatever the computational cost, then we recommend doing multiple sample realignment so that novel indels in one sample help to realign reads in other samples.  Although not generally necessary for deep sequencing data, this is important for low-coverage multi-sample SNP calling projects like the 1000 Genomes Project.  Note that the computational cost here is so extreme that we only do this analysis in special circumstances, such as large-scale data freeze for the project.  &lt;br /&gt;
&lt;br /&gt;
Note that for contrastive calling projects -- such as cancer tumor/normals -- that we recommend cleaning both the tumor and the normal together in general to avoid slight alignment differences between the two tissue types.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each sample&lt;br /&gt;
    lanes.bam &amp;lt;- merged lane.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(lanes.bam)&lt;br /&gt;
&lt;br /&gt;
samples.bam &amp;lt;- merged dedup.bam&#039;s for all samples&lt;br /&gt;
realigned.bam &amp;lt;- realign(samples.bam)&lt;br /&gt;
recal.bam &amp;lt;- recal(realigned.bam)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Misc. notes on the process ====&lt;br /&gt;
&lt;br /&gt;
* MarkDuplicates needs only be run at the library level. So the sample-level dedupping isn&#039;t necessary if you only ever a library on a single lane.  If you run the sample library on many lanes (as can be necessary for whole exome, for example), you should dedup at the library level.&lt;br /&gt;
* The base quality score recalibrator is read group aware, so running it on a merged BAM files containing multiple read groups is the same as running it on each bam file individually.  There&#039;s some memory cost (so it&#039;s best not to recalibrate 10000 RGs simultaneously) but for reasonable projects this is a fine.&lt;br /&gt;
* Local realignment preserves read meta-data, so you can realign and then recalibrate just fine.&lt;br /&gt;
&lt;br /&gt;
== Initial variant discovery and genotyping ==&lt;br /&gt;
&lt;br /&gt;
=== Input BAMs for variant discovery and genotyping ===&lt;br /&gt;
&lt;br /&gt;
After the raw data processing step, the GATK variant detection process assumes that you have aligned, duplicate marked, and recalibrated BAM files for all of the samples in your cohort.  Because the GATK can dynamically merge BAM files, it isn&#039;t critical to have merged files by lane into sample bams, or even samples bams into cohort bams.  In general we try to create sample level bams for deep data sets (deep WG or exomes) and merged cohort files by chromosome for WG low-pass.  A nice size for BAMs is 10-300 Gb, just for organizing on disk.  &lt;br /&gt;
&lt;br /&gt;
For this part of the this document, I&#039;m going to assume that you have a single realigned, recalibrated, dedupped BAM per sample, called sampleX.bam, for X from 1 to N samples in your cohort.  Note that some of the data processing steps, such as multiple sample local realignment, will merge BAMS for many samples into a single BAM.  If you&#039;ve gone down this route, you just need to modify the GATK commands as necessary to take not multiple BAMs, one for each sample, but a single BAM for all samples.&lt;br /&gt;
&lt;br /&gt;
=== Multi-sample SNP and indel calling ===&lt;br /&gt;
&lt;br /&gt;
The next step in the standard GATK data processing pipeline, whole genome or targeted, deep or shallow, is to apply the [[Unified genotyper]] to identify sites among the cohort samples that are statistically non-reference.  This will produce a multi-sample [[VCF Format|VCF file]], with sites discovered across samples and genotypes assigned to each sample in the cohort.  It&#039;s in this stage that we use the meta-data in the BAM files extensively -- read groups for reads, with samples, platforms, etc -- to enable us to do the multi-sample merging and genotyping correctly.  It was a pain for data processing, yes, but now life is easy for downstream calling and analysis.&lt;br /&gt;
&lt;br /&gt;
Note that by default the Unified Genotyper calls SNPs only.  To enable the indel calling capabilities instead use the -glm DINDEL argument.&lt;br /&gt;
&lt;br /&gt;
==== Selecting an appropriate quality score threshold ====&lt;br /&gt;
&lt;br /&gt;
A common question is the confidence score threshold to use for variant detection.  We recommend:&lt;br /&gt;
&lt;br /&gt;
; Deep (&amp;gt; 10x coverage per sample) data : we recommend a minimum confidence score threshold of Q30 with an emission threshold of Q10.  These Q10-Q30 calls will be emitted filtered out as LowQual.  &lt;br /&gt;
; Shallow (&amp;lt; 10x coverage per sample) data : because variants have by necessity lower quality with shallower coverage, we recommend a min. confidence score of Q4 and an emission threshold of Q3. &lt;br /&gt;
&lt;br /&gt;
If you want include these LowQual variants following [[Variant quality score recalibration]] you can just ignore the LowQual filter in the VariantRecalibration step.  That way they won&#039;t influence the recalibration but can be selected, if they appear sufficiently good, for inclusion in the final call set.&lt;br /&gt;
&lt;br /&gt;
=== Protocol ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
raw.vcf &amp;lt;- unifiedGenotyper(sample1.bam, sample2.bam, ..., sampleN.bam)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Integrating analyses: getting the best call set possible ==&lt;br /&gt;
&lt;br /&gt;
This raw VCF file should be as sensitive to variation as you&#039;ll get without imputation.  At this stage, you can assess things like sensitivity to known variant sites or genotype chip concordance.  The problem is that the raw VCF will have many sites that aren&#039;t really genetic variants but are machine artifacts that make the site statistically non-reference.  All of the subsequent filtering and/or quality score recalibration steps are designed to separate out the FP machine artifacts from the TP genetic variants.&lt;br /&gt;
&lt;br /&gt;
The tools used here is [[VariantFiltrationWalker]] to apply hard filters and [[Variant quality score recalibration]] to build an adaptive error model using known variant sites and then apply this model to estimate the probability that each variant is a true genetic variant or a machine artifact.  Regardless of whether you&#039;ll ultimately apply hard filtering or adaptive error modeling to select your final calls, we recommend you first apply some common SNP filters to avoid obvious misalignment and indel artifacts.  &lt;br /&gt;
&lt;br /&gt;
=== Analysis read VCF procotol ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
snpFiltered.vcf &amp;lt;- basicSNPFilters(raw.vcf)&lt;br /&gt;
snpFiltered.indelFiltered.vcf &amp;lt;- indelFilters(snpFiltered.vcf)&lt;br /&gt;
&lt;br /&gt;
if we are hard filtering:&lt;br /&gt;
    analysisReady.vcf &amp;lt;- dataTypeSpecificFilters(snpFiltered.indelFiltered.vcf)&lt;br /&gt;
&lt;br /&gt;
else we are doing variant quality score recalibration&lt;br /&gt;
    analysisReady.vcf &amp;lt;- variantQualityScoreRecalibrate(snpFiltered.indelFiltered.vcf)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note that [[VariantFiltrationWalker]] can likely apply simultaneously all of the filters you want.  For easy of explanation, we divide up the application of these filters into independent applications of the [[VariantFiltrationWalker]]. Overall, the approximate command for variant filtering will look something like, but depending on exactly what how you want to filter your variants the detail of this command will change:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Common VariantFiltrationWalker command&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
java -jar GenomeAnalysisTK.jar \&lt;br /&gt;
  -T VariantFiltration \&lt;br /&gt;
  -R resources/Homo_sapiens_assembly18.fasta \&lt;br /&gt;
  -B:variant,VCF input.vcf \&lt;br /&gt;
  -o filtered.vcf \&lt;br /&gt;
  [FILTERING PARAMETERS]&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Basic indel filtering ===&lt;br /&gt;
&lt;br /&gt;
The goal of the Basic indel filters is to remove alignment artifacts from the data.  The key item here is flagging variants with high strand bias and in poorly mapped regions (HARD_TO_VALIDATE set) with more than 10% of the reads having mapping quality 0.&lt;br /&gt;
Arguments for [[VariantFiltrationWalker]]:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  --filterExpression &amp;quot;MQ0 &amp;gt;= 4 &amp;amp;&amp;amp; ((MQ0 / (1.0 * DP)) &amp;gt; 0.1)&amp;quot;    [expression to match 10% of reads with MAPQ0]&lt;br /&gt;
  --filterName &amp;quot;HARD_TO_VALIDATE&amp;quot;&lt;br /&gt;
  --filterExpression &amp;quot;SB &amp;gt;= -1.0&amp;quot;&lt;br /&gt;
  --filterName &amp;quot;StrandBiasFilter&amp;quot;&lt;br /&gt;
  --filterExpression &amp;quot;QUAL &amp;lt; 10&amp;quot;&lt;br /&gt;
  --filterName &amp;quot;QualFilter&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note that, for the above to work, the input vcf needs to be annotated with the corresponding values (MQ0,DP,SB). If any of these values are missing, then VariantAnnotator needs to be run first so that VariantFiltration can run properly.&lt;br /&gt;
&lt;br /&gt;
=== Basic SNP filtering ===&lt;br /&gt;
&lt;br /&gt;
The goal of the Basic SNP filters is to remove alignment artifacts from the data.  The key items here is flagging SNPs within clusters (3 SNPs with 10 bp of each other) and those in poorly mapped regions (HARD_TO_VALIDATE set) with more than 10% of the reads having mapping quality 0.&lt;br /&gt;
&lt;br /&gt;
Arguments for [[VariantFiltrationWalker]]:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  --clusterWindowSize 10 &lt;br /&gt;
  --filterExpression &amp;quot;MQ0 &amp;gt;= 4 &amp;amp;&amp;amp; ((MQ0 / (1.0 * DP)) &amp;gt; 0.1)&amp;quot;    [expression to match 10% of reads with MAPQ0]&lt;br /&gt;
  --filterName &amp;quot;HARD_TO_VALIDATE&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The clusterWindowSize can be reduced below 10 if you expect many true variants close together, such as in an organisms with high \math{theta} or with many individuals (1000s of humans).&lt;br /&gt;
&lt;br /&gt;
As mentioned above, this command assumes that MQ0 and DP are present in the input vcf. If they are not, then you need to run VariantAnnotator first to complete the corresponding annotations.&lt;br /&gt;
&lt;br /&gt;
=== Filtering around indels ===&lt;br /&gt;
&lt;br /&gt;
It&#039;s possible that, despite even local realignment, misalignments around true and artifactual indels will result in some false SNP calls.  These errors are quite common if you didn&#039;t do local realignment, didn&#039;t provide a set of known indels during local realignment, and around very large indels that can&#039;t be modeled properly by local realignment.  If you performed indel calling then you can filter your SNP calls around the raw indel calls from your data set.  The raw calls should be used so that you can remove SNPs around even artifactual indels, which even though the indel caller doesn&#039;t believe they are real, are sufficient to generate FP SNP calls.  Skipping this step will leave some variants in your call set that would have been identified as indel artifacts.  It&#039;s possible that downstream filters or variant quality score recalibration will remove these, but if you want the best calls possible, it&#039;s worth the effort to indel call for the mask alone.&lt;br /&gt;
&lt;br /&gt;
Arguments for [[VariantFiltrationWalker]]:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  -B:mask,VCF /path/to/indels.from.UnifiedGenotyper.vcf&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
If you have good indel calls for your population (for example, the 1000 Genomes Project pilot indel calls that include 1.3M sites) you can filter calls around these indels by generating a mask for those indels.  This is a more advanced approach that&#039;s only really necessary if you don&#039;t realign around indels, but can help avoid even more indel artifacts.  If these indels are not available in VCF format, you can create a [http://genome.ucsc.edu/FAQ/FAQformat.html#format1 Bed] file with the loci to mask, and supply the mask with the following command:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  -B:mask,Bed /path/to/indels.mask.bed&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Recall that the bed format is 0-based and half-closed.&lt;br /&gt;
&lt;br /&gt;
; Deep (&amp;gt; 10x coverage per sample) data : we recommend the raw indel calls from UnifiedGenotyper be supplied as the mask.&lt;br /&gt;
; Shallow (&amp;lt; 10x coverage per sample) data : given the soft-calling procedure, which can learn the appropriate filters automatically, an indel mask is not required.&lt;br /&gt;
&lt;br /&gt;
=== Making analysis ready calls SNP calls with hard filtering ===&lt;br /&gt;
&lt;br /&gt;
We include these hard filter parameters for several reasons.  One, they give reasonably good results with sufficient depth and multiple samples.  Second, many data sets were generated and filtered with these parameters and we want to make sure that people can generate similar data sets in the future.  Finally, VQSR requires lots of data to work properly, and with targeted resequencing of a small region (for example, a few hundred genes) VQSR may be inappropriate leaving hard filtering as the only option.  &lt;br /&gt;
&lt;br /&gt;
Arguments for [[VariantFiltrationWalker]]:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  --filterExpression [DATA_TYPE_SPECIFIC_FILTERS]&lt;br /&gt;
  --filterName GATKStandard&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where DATA_TYPE_SPECIFIC_FILTERS has project-specific filtering strings selected below:&lt;br /&gt;
&lt;br /&gt;
* For exomes with deep coverage per sample &lt;br /&gt;
** DATA_TYPE_SPECIFIC_FILTERS should be &amp;quot;QUAL &amp;lt; 30.0 || QD &amp;lt; 5.0 || HRun &amp;gt; 5 || SB &amp;gt; -0.10&amp;quot;&lt;br /&gt;
&lt;br /&gt;
* Whole genomes with deep coverage:  &lt;br /&gt;
** DATA_TYPE_SPECIFIC_FILTERS should be &amp;quot;DP &amp;gt; 100 || MQ0 &amp;gt; 40 || SB &amp;gt; -0.10&amp;quot;. &lt;br /&gt;
** Note that the DP values (depth filters) must be parameterized by the average coverage among all of the samples.  That is, if you have 100x mean coverage, there&#039;s nothing wrong with a site with 110x, but usually 1000x indicates a bad sites, at least whole genome.&lt;br /&gt;
&lt;br /&gt;
* Shallow-coverage (&amp;lt;10x) : you cannot use filtering to reliably separate TPs from FPs.  You must use the protocol involving [[Variant quality score recalibration]]&lt;br /&gt;
&lt;br /&gt;
The maximum DP (depth) filter only applies to whole genome data, where the probability of a site having exactly N reads given an average coverage of M is a well-behaved function.  First principles suggest this should be a binomial sampling but in practice it is more a Gaussian distribution.  Regardless, the DP threshold should be set a 5 or 6 sigma from the mean coverage across all samples, so that the DP &amp;gt; X threshold eliminates sites with excessive coverage caused by alignment artifacts.  Note that &#039;&#039;&#039;For exomes, a straight DP filter shouldn&#039;t be used&#039;&#039;&#039; because the relationship between misalignments and depth isn&#039;t clear for capture data.  Note that as of Dec. 2010 we no longer recommend using AB to filter, as (1) it&#039;s not generated by the UG by default any more and (2) QD is a better, and largely equivalent, filter.&lt;br /&gt;
&lt;br /&gt;
That said, all of the caveats about determining the right parameters, etc, are annoying and are largely eliminated by [[Variant quality score recalibration]].&lt;br /&gt;
&lt;br /&gt;
=== Making analysis ready calls with variant quality score recalibration ===&lt;br /&gt;
&lt;br /&gt;
An alternative approach to hard filtering, developed over the last 6 months in GSA to handle the low-pass whole genome data sets from the 1000 Genomes Project is an approach called [[Variant quality score recalibration]], described in detail on that wiki.  Guidelines for applying the [[Variant quality score recalibration]] are listed on the wiki, including which annotations to use and how to slice the recalibrated call set into the high-quality call sets.  The input call set to [[Variant quality score recalibration]] should be the snpFiltered.indelFiltered.vcf.  Note that the HARD_TO_VALIDATE filter are respected during training of the error model but these SNPs will be scored a potentially true variants, as this approach appears to robustly differentiate high-quality calls in the HARD_TO_VALIDATE set from the unreliable ones.&lt;br /&gt;
&lt;br /&gt;
== Expected SNP call quality ==&lt;br /&gt;
&lt;br /&gt;
All of the following tables were generated using [[VariantEval]].  You can generate your own data points for comparing with that tool and the correct input / comparison VCF files.&lt;br /&gt;
&lt;br /&gt;
=== Summary results for deep whole genome, multi-sample low-pass, and whole exome ===&lt;br /&gt;
&lt;br /&gt;
[[File:best_practice_calls.jpg|400px|thumb|Summary results for best practice calling pipeline]]&lt;br /&gt;
&lt;br /&gt;
The associated table provides some expectations for running deep whole genomes, single whole exomes, and multi-sample low-pass (from the 1000 genomes) in terms of sensitivity, specificity, and Ti/Tv ratios for known and novel calls.  Obviously individual data sets will be different but this demonstrates that running the pipeline described here ( realigned -&amp;gt; recalibration -&amp;gt; calling -&amp;gt; filtering -&amp;gt; variant quality score recalibration) can produce excellent results for a variety of data types.  Note that the deep data sets are single sample; in our hands multi-sample deep data results in much better calls overall.&lt;br /&gt;
&lt;br /&gt;
=== Expected Ti/Tv ratios ===&lt;br /&gt;
&lt;br /&gt;
[[File:titv_expectations.jpg|400px|thumb|Expected known and novel Ti/Tv ratios]]&lt;br /&gt;
&lt;br /&gt;
We provide two useful data points for evaluating the quality of SNP calls whole genome or in the targeted whole exome (Agilent).  You can follow a similar methodology to establish the Ti/Tv expectations for your region of interest, if captured separately, by running [[VariantEval]] on the 1000 Genomes Trios or Complete Genomes or other highly reliable data sets given the targeted interval.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Previous versions of Best Practices (now outdated) ==&lt;br /&gt;
===Version 1 -- hard filters===&lt;br /&gt;
* [[Whole exome v1]]&lt;br /&gt;
* [[Whole genome, deep coverage v1]]&lt;br /&gt;
* [[Whole genome, low-pass v1]]&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Sandbox:Best_Practice_Variant_Detection_with_the_GATK_v2&amp;diff=95006</id>
		<title>Sandbox:Best Practice Variant Detection with the GATK v2</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Sandbox:Best_Practice_Variant_Detection_with_the_GATK_v2&amp;diff=95006"/>
		<updated>2011-05-18T16:41:26Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Data Processing Pipeline Script ==&lt;br /&gt;
The [[Data Processing Pipeline]] is a Queue script that performs all data processing described in this page following our most current recommendations for best practices.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
Our current best practice for making SNP calls is divided into 5 sequential steps (see Figure): initial mapping, refinement of the initial reads, multi-sample indel and SNP calling, filtering of the raw SNP calls, and finally variant quality score recalibration.  These steps are the same for targeted resequencing, whole exomes, deep whole genomes, and low-pass whole genomes.  The exact commands for each tool are available on the tools wiki entry.   &lt;br /&gt;
&lt;br /&gt;
Note that, due to the specific attributes of a project, that the the specific values used in each of the commands may need to be selected by the analyst.  Care should be taken by the analyst running our tools to understand what each parameter does and to evaluate which value best fits his/her data.&lt;br /&gt;
&lt;br /&gt;
[[File:http://www.broadinstitute.org/gsa/wiki/images/7/7a/overall_flow.jpg|300px|thumb|Overview of GATK data processing pipeline]]&lt;br /&gt;
&lt;br /&gt;
=== Lanes, Samples, Cohort ===&lt;br /&gt;
&lt;br /&gt;
There are four major organizational units for next-generation DNA sequencing processes:&lt;br /&gt;
&lt;br /&gt;
; Lane : The basic machine unit for sequencing.  The lane reflects the basic independent run of an NGS machine.  For Illumina machines, this is the physical sequencing lane.  &lt;br /&gt;
&lt;br /&gt;
; Library : A unit of DNA preparation that at some point is physically pooled together.  Multiple lanes can be run from aliquots from the same library.  The DNA library and its preparation is the natural unit that is being sequenced.  For example, if the library has limited complexity, then many sequences are duplicated and will result in a high duplication rate across lanes. &lt;br /&gt;
&lt;br /&gt;
; Sample : A single individual, such as human CEPH NA12878.  Multiple libraries with different properties can be constructed from the original sample DNA source.  Here we treat samples as independent individuals whose genome sequence we are attempting to determine.  From this perspective, tumor / normal samples are different despite coming from the same individual.&lt;br /&gt;
&lt;br /&gt;
; Cohort : A collection of samples being analyzed together.  This organizational unit is the most subjective and depends intimately on the design goals of the sequencing project.  For population discovery projects like the 1000 Genomes, the analysis cohort is the ~100 individual in each population.  For exome projects with many samples (e.g., ESP with 800 EOMI samples) deeply sequenced we divide up the complete set of samples into cohorts of ~50 individuals for multi-sample analyses.  &lt;br /&gt;
&lt;br /&gt;
This document describes how to call variation within a single analysis cohort, comprised for one or many samples, each of one or many libraries that were sequenced on at least one lane of an NGS machine. &lt;br /&gt;
&lt;br /&gt;
Note that many GATK commands can be run at the lane level, but will give better results seeing all of the data for a single sample, or even all of the data for all samples.  Unfortunately, there&#039;s a trade-off in computational cost by running these commands across all of your data simultaneously.&lt;br /&gt;
&lt;br /&gt;
=== Testing data: 64x HiSeq on chr20 for NA12878 ===&lt;br /&gt;
&lt;br /&gt;
In order to help individuals get up to speed, evaluate their command lines, and generally become familiar with the GATK tools we recommend you download the raw and realigned, recalibrated [[NA12878 test data]].  It should be possible to apply all of the approaches outlined below to get excellent results for realignment, recalibration, SNP calling, indel calling, filtering, and variant quality score recalibration using this data.&lt;br /&gt;
&lt;br /&gt;
== Phase I: Raw data processing ==&lt;br /&gt;
[[Image:DataProcessingPPL.jpg|thumb|300px|Data processing pipeline of the best practices for raw data processing, from sequencer data (fastq files) to analysis read reads (bam file)]]&lt;br /&gt;
&lt;br /&gt;
=== Initial read mapping ===&lt;br /&gt;
&lt;br /&gt;
The GATK data processing pipeline assumes that one of the many NGS read aligners (see [http://bib.oxfordjournals.org/content/11/5/473.abstract] for a review) has been applied to your raw FASTQ files.  For Illumina data we recommend [http://bio-bwa.sourceforge.net/ BWA] because it is accurate, fast, well-supported, open-source, and emits BAM files natively.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Raw BAM to realigned, recalibrated BAM ===&lt;br /&gt;
&lt;br /&gt;
The three key tools here are [[Base quality score recalibration]], [[Local realignment around indels]], and [http://picard.sourceforge.net/command-line-overview.shtml#MarkDuplicates|Picard&#039;s MarkDuplicates].  Although ideally one follows the flow from Figure 1, in practice MarkDuplicates can be run before local realignment, in order to handle cases where duplicates overlapping indels get marginally different alignments   (unlikely but possible) and so will not be considered as potential duplicates because MarkDuplicates looks only at read pair start and stop positions.  There are several options here, from the easy and fast basic protocol to the more.&lt;br /&gt;
&lt;br /&gt;
There are two types of realignment:&lt;br /&gt;
* Realignment only at known sites, which is very efficient, can operate with little coverage (1x per lane genome wide) but can only realign reads at known indels.&lt;br /&gt;
* Fully local realignment uses mismatching bases to determine if a site should be realigned, and relies on sufficient coverage to discover the correct indel allele in the reads for alignment.  It is much slower (involves SW step) but can discover new indel sites in the reads.  If you have a database of known indels (for human, this database is extensive) then at this stage you would also include these indels during realignment, which vastly improves sensitivity, specificity, and speed.&lt;br /&gt;
&lt;br /&gt;
==== Previous recommendation: lane-level recalibration, sample-level realignment ====&lt;br /&gt;
&lt;br /&gt;
This is the protocol we used at the Broad Institute for the last year.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each lane.bam&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicate(lane.bam)&lt;br /&gt;
    recal.bam &amp;lt;- recal(dedup.bam)&lt;br /&gt;
&lt;br /&gt;
for each sample&lt;br /&gt;
    lanes.bam &amp;lt;- merged recal.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(lanes.bam)&lt;br /&gt;
    realigned.bam &amp;lt;- realign(dedup.bam) [with known sites if possible]&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fast: lane-level realignment at known sites only and lane-level recalibration ====&lt;br /&gt;
&lt;br /&gt;
This protocol adds lane-level local realignment around known indels, making it very fast (no sample level processing) and gives better results for human samples than the previous recommendation:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each lane.bam&lt;br /&gt;
    realigned.bam &amp;lt;- realign(lane.bam) [at only known sites]&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicate(realigned.bam)&lt;br /&gt;
    recal.bam &amp;lt;- recal(dedup.bam)&lt;br /&gt;
&lt;br /&gt;
for each sample&lt;br /&gt;
    recals.bam &amp;lt;- merged lane-level recal.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(recals.bam)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fast + sample-level realignment ====&lt;br /&gt;
&lt;br /&gt;
This protocol adds sample-level realignment after library / sample level dedupping, so that novel indels in each sample can be discovered and realigned around.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each lane.bam&lt;br /&gt;
    realigned.bam &amp;lt;- realign(lane.bam) [at only known sites]&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicate(realigned.bam)&lt;br /&gt;
    recal.bam &amp;lt;- recal(dedup.bam)&lt;br /&gt;
&lt;br /&gt;
for each sample&lt;br /&gt;
    recals.bam &amp;lt;- merged lane-level recal.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(recals.bam)&lt;br /&gt;
    realigned.bam &amp;lt;- realign(dedup.bam) [with known sites if possible]&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Better: sample-level realignment with known indels and recalibration ====&lt;br /&gt;
&lt;br /&gt;
Rather than doing the lane level cleaning and recalibration, this process aggregates all of the reads for each sample and then does a full dedupping, realign, and recalibration, yielding the best single-sample results.  The big change here is sample-level cleaning followed by recalibration, giving you the most accurate quality scores possible for a single sample. &lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each sample&lt;br /&gt;
    lanes.bam &amp;lt;- merged lane.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(lanes.bam)&lt;br /&gt;
    realigned.bam &amp;lt;- realign(dedup.bam) [with known sites if possible]&lt;br /&gt;
    recal.bam &amp;lt;- recal(realigned.bam)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Best: multi-sample realignment with known sites and recalibration ====&lt;br /&gt;
&lt;br /&gt;
Finally, if you really want to get the absolute best results, whatever the computational cost, then we recommend doing multiple sample realignment so that novel indels in one sample help to realign reads in other samples.  Although not generally necessary for deep sequencing data, this is important for low-coverage multi-sample SNP calling projects like the 1000 Genomes Project.  Note that the computational cost here is so extreme that we only do this analysis in special circumstances, such as large-scale data freeze for the project.  &lt;br /&gt;
&lt;br /&gt;
Note that for contrastive calling projects -- such as cancer tumor/normals -- that we recommend cleaning both the tumor and the normal together in general to avoid slight alignment differences between the two tissue types.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each sample&lt;br /&gt;
    lanes.bam &amp;lt;- merged lane.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(lanes.bam)&lt;br /&gt;
&lt;br /&gt;
samples.bam &amp;lt;- merged dedup.bam&#039;s for all samples&lt;br /&gt;
realigned.bam &amp;lt;- realign(samples.bam)&lt;br /&gt;
recal.bam &amp;lt;- recal(realigned.bam)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Misc. notes on the process ====&lt;br /&gt;
&lt;br /&gt;
* MarkDuplicates needs only be run at the library level. So the sample-level dedupping isn&#039;t necessary if you only ever a library on a single lane.  If you run the sample library on many lanes (as can be necessary for whole exome, for example), you should dedup at the library level.&lt;br /&gt;
* The base quality score recalibrator is read group aware, so running it on a merged BAM files containing multiple read groups is the same as running it on each bam file individually.  There&#039;s some memory cost (so it&#039;s best not to recalibrate 10000 RGs simultaneously) but for reasonable projects this is a fine.&lt;br /&gt;
* Local realignment preserves read meta-data, so you can realign and then recalibrate just fine.&lt;br /&gt;
&lt;br /&gt;
== Initial variant discovery and genotyping ==&lt;br /&gt;
&lt;br /&gt;
=== Input BAMs for variant discovery and genotyping ===&lt;br /&gt;
&lt;br /&gt;
After the raw data processing step, the GATK variant detection process assumes that you have aligned, duplicate marked, and recalibrated BAM files for all of the samples in your cohort.  Because the GATK can dynamically merge BAM files, it isn&#039;t critical to have merged files by lane into sample bams, or even samples bams into cohort bams.  In general we try to create sample level bams for deep data sets (deep WG or exomes) and merged cohort files by chromosome for WG low-pass.  A nice size for BAMs is 10-300 Gb, just for organizing on disk.  &lt;br /&gt;
&lt;br /&gt;
For this part of the this document, I&#039;m going to assume that you have a single realigned, recalibrated, dedupped BAM per sample, called sampleX.bam, for X from 1 to N samples in your cohort.  Note that some of the data processing steps, such as multiple sample local realignment, will merge BAMS for many samples into a single BAM.  If you&#039;ve gone down this route, you just need to modify the GATK commands as necessary to take not multiple BAMs, one for each sample, but a single BAM for all samples.&lt;br /&gt;
&lt;br /&gt;
=== Multi-sample SNP and indel calling ===&lt;br /&gt;
&lt;br /&gt;
The next step in the standard GATK data processing pipeline, whole genome or targeted, deep or shallow, is to apply the [[Unified genotyper]] to identify sites among the cohort samples that are statistically non-reference.  This will produce a multi-sample [[VCF Format|VCF file]], with sites discovered across samples and genotypes assigned to each sample in the cohort.  It&#039;s in this stage that we use the meta-data in the BAM files extensively -- read groups for reads, with samples, platforms, etc -- to enable us to do the multi-sample merging and genotyping correctly.  It was a pain for data processing, yes, but now life is easy for downstream calling and analysis.&lt;br /&gt;
&lt;br /&gt;
Note that by default the Unified Genotyper calls SNPs only.  To enable the indel calling capabilities instead use the -glm DINDEL argument.&lt;br /&gt;
&lt;br /&gt;
==== Selecting an appropriate quality score threshold ====&lt;br /&gt;
&lt;br /&gt;
A common question is the confidence score threshold to use for variant detection.  We recommend:&lt;br /&gt;
&lt;br /&gt;
; Deep (&amp;gt; 10x coverage per sample) data : we recommend a minimum confidence score threshold of Q30 with an emission threshold of Q10.  These Q10-Q30 calls will be emitted filtered out as LowQual.  &lt;br /&gt;
; Shallow (&amp;lt; 10x coverage per sample) data : because variants have by necessity lower quality with shallower coverage, we recommend a min. confidence score of Q4 and an emission threshold of Q3. &lt;br /&gt;
&lt;br /&gt;
If you want include these LowQual variants following [[Variant quality score recalibration]] you can just ignore the LowQual filter in the VariantRecalibration step.  That way they won&#039;t influence the recalibration but can be selected, if they appear sufficiently good, for inclusion in the final call set.&lt;br /&gt;
&lt;br /&gt;
=== Protocol ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
raw.vcf &amp;lt;- unifiedGenotyper(sample1.bam, sample2.bam, ..., sampleN.bam)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Integrating analyses: getting the best call set possible ==&lt;br /&gt;
&lt;br /&gt;
This raw VCF file should be as sensitive to variation as you&#039;ll get without imputation.  At this stage, you can assess things like sensitivity to known variant sites or genotype chip concordance.  The problem is that the raw VCF will have many sites that aren&#039;t really genetic variants but are machine artifacts that make the site statistically non-reference.  All of the subsequent filtering and/or quality score recalibration steps are designed to separate out the FP machine artifacts from the TP genetic variants.&lt;br /&gt;
&lt;br /&gt;
The tools used here is [[VariantFiltrationWalker]] to apply hard filters and [[Variant quality score recalibration]] to build an adaptive error model using known variant sites and then apply this model to estimate the probability that each variant is a true genetic variant or a machine artifact.  Regardless of whether you&#039;ll ultimately apply hard filtering or adaptive error modeling to select your final calls, we recommend you first apply some common SNP filters to avoid obvious misalignment and indel artifacts.  &lt;br /&gt;
&lt;br /&gt;
=== Analysis read VCF procotol ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
snpFiltered.vcf &amp;lt;- basicSNPFilters(raw.vcf)&lt;br /&gt;
snpFiltered.indelFiltered.vcf &amp;lt;- indelFilters(snpFiltered.vcf)&lt;br /&gt;
&lt;br /&gt;
if we are hard filtering:&lt;br /&gt;
    analysisReady.vcf &amp;lt;- dataTypeSpecificFilters(snpFiltered.indelFiltered.vcf)&lt;br /&gt;
&lt;br /&gt;
else we are doing variant quality score recalibration&lt;br /&gt;
    analysisReady.vcf &amp;lt;- variantQualityScoreRecalibrate(snpFiltered.indelFiltered.vcf)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note that [[VariantFiltrationWalker]] can likely apply simultaneously all of the filters you want.  For easy of explanation, we divide up the application of these filters into independent applications of the [[VariantFiltrationWalker]]. Overall, the approximate command for variant filtering will look something like, but depending on exactly what how you want to filter your variants the detail of this command will change:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Common VariantFiltrationWalker command&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
java -jar GenomeAnalysisTK.jar \&lt;br /&gt;
  -T VariantFiltration \&lt;br /&gt;
  -R resources/Homo_sapiens_assembly18.fasta \&lt;br /&gt;
  -B:variant,VCF input.vcf \&lt;br /&gt;
  -o filtered.vcf \&lt;br /&gt;
  [FILTERING PARAMETERS]&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Basic indel filtering ===&lt;br /&gt;
&lt;br /&gt;
The goal of the Basic indel filters is to remove alignment artifacts from the data.  The key item here is flagging variants with high strand bias and in poorly mapped regions (HARD_TO_VALIDATE set) with more than 10% of the reads having mapping quality 0.&lt;br /&gt;
Arguments for [[VariantFiltrationWalker]]:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  --filterExpression &amp;quot;MQ0 &amp;gt;= 4 &amp;amp;&amp;amp; ((MQ0 / (1.0 * DP)) &amp;gt; 0.1)&amp;quot;    [expression to match 10% of reads with MAPQ0]&lt;br /&gt;
  --filterName &amp;quot;HARD_TO_VALIDATE&amp;quot;&lt;br /&gt;
  --filterExpression &amp;quot;SB &amp;gt;= -1.0&amp;quot;&lt;br /&gt;
  --filterName &amp;quot;StrandBiasFilter&amp;quot;&lt;br /&gt;
  --filterExpression &amp;quot;QUAL &amp;lt; 10&amp;quot;&lt;br /&gt;
  --filterName &amp;quot;QualFilter&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note that, for the above to work, the input vcf needs to be annotated with the corresponding values (MQ0,DP,SB). If any of these values are missing, then VariantAnnotator needs to be run first so that VariantFiltration can run properly.&lt;br /&gt;
&lt;br /&gt;
=== Basic SNP filtering ===&lt;br /&gt;
&lt;br /&gt;
The goal of the Basic SNP filters is to remove alignment artifacts from the data.  The key items here is flagging SNPs within clusters (3 SNPs with 10 bp of each other) and those in poorly mapped regions (HARD_TO_VALIDATE set) with more than 10% of the reads having mapping quality 0.&lt;br /&gt;
&lt;br /&gt;
Arguments for [[VariantFiltrationWalker]]:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  --clusterWindowSize 10 &lt;br /&gt;
  --filterExpression &amp;quot;MQ0 &amp;gt;= 4 &amp;amp;&amp;amp; ((MQ0 / (1.0 * DP)) &amp;gt; 0.1)&amp;quot;    [expression to match 10% of reads with MAPQ0]&lt;br /&gt;
  --filterName &amp;quot;HARD_TO_VALIDATE&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The clusterWindowSize can be reduced below 10 if you expect many true variants close together, such as in an organisms with high \math{theta} or with many individuals (1000s of humans).&lt;br /&gt;
&lt;br /&gt;
As mentioned above, this command assumes that MQ0 and DP are present in the input vcf. If they are not, then you need to run VariantAnnotator first to complete the corresponding annotations.&lt;br /&gt;
&lt;br /&gt;
=== Filtering around indels ===&lt;br /&gt;
&lt;br /&gt;
It&#039;s possible that, despite even local realignment, misalignments around true and artifactual indels will result in some false SNP calls.  These errors are quite common if you didn&#039;t do local realignment, didn&#039;t provide a set of known indels during local realignment, and around very large indels that can&#039;t be modeled properly by local realignment.  If you performed indel calling then you can filter your SNP calls around the raw indel calls from your data set.  The raw calls should be used so that you can remove SNPs around even artifactual indels, which even though the indel caller doesn&#039;t believe they are real, are sufficient to generate FP SNP calls.  Skipping this step will leave some variants in your call set that would have been identified as indel artifacts.  It&#039;s possible that downstream filters or variant quality score recalibration will remove these, but if you want the best calls possible, it&#039;s worth the effort to indel call for the mask alone.&lt;br /&gt;
&lt;br /&gt;
Arguments for [[VariantFiltrationWalker]]:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  -B:mask,VCF /path/to/indels.from.UnifiedGenotyper.vcf&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
If you have good indel calls for your population (for example, the 1000 Genomes Project pilot indel calls that include 1.3M sites) you can filter calls around these indels by generating a mask for those indels.  This is a more advanced approach that&#039;s only really necessary if you don&#039;t realign around indels, but can help avoid even more indel artifacts.  If these indels are not available in VCF format, you can create a [http://genome.ucsc.edu/FAQ/FAQformat.html#format1 Bed] file with the loci to mask, and supply the mask with the following command:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  -B:mask,Bed /path/to/indels.mask.bed&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Recall that the bed format is 0-based and half-closed.&lt;br /&gt;
&lt;br /&gt;
; Deep (&amp;gt; 10x coverage per sample) data : we recommend the raw indel calls from UnifiedGenotyper be supplied as the mask.&lt;br /&gt;
; Shallow (&amp;lt; 10x coverage per sample) data : given the soft-calling procedure, which can learn the appropriate filters automatically, an indel mask is not required.&lt;br /&gt;
&lt;br /&gt;
=== Making analysis ready calls SNP calls with hard filtering ===&lt;br /&gt;
&lt;br /&gt;
We include these hard filter parameters for several reasons.  One, they give reasonably good results with sufficient depth and multiple samples.  Second, many data sets were generated and filtered with these parameters and we want to make sure that people can generate similar data sets in the future.  Finally, VQSR requires lots of data to work properly, and with targeted resequencing of a small region (for example, a few hundred genes) VQSR may be inappropriate leaving hard filtering as the only option.  &lt;br /&gt;
&lt;br /&gt;
Arguments for [[VariantFiltrationWalker]]:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  --filterExpression [DATA_TYPE_SPECIFIC_FILTERS]&lt;br /&gt;
  --filterName GATKStandard&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where DATA_TYPE_SPECIFIC_FILTERS has project-specific filtering strings selected below:&lt;br /&gt;
&lt;br /&gt;
* For exomes with deep coverage per sample &lt;br /&gt;
** DATA_TYPE_SPECIFIC_FILTERS should be &amp;quot;QUAL &amp;lt; 30.0 || QD &amp;lt; 5.0 || HRun &amp;gt; 5 || SB &amp;gt; -0.10&amp;quot;&lt;br /&gt;
&lt;br /&gt;
* Whole genomes with deep coverage:  &lt;br /&gt;
** DATA_TYPE_SPECIFIC_FILTERS should be &amp;quot;DP &amp;gt; 100 || MQ0 &amp;gt; 40 || SB &amp;gt; -0.10&amp;quot;. &lt;br /&gt;
** Note that the DP values (depth filters) must be parameterized by the average coverage among all of the samples.  That is, if you have 100x mean coverage, there&#039;s nothing wrong with a site with 110x, but usually 1000x indicates a bad sites, at least whole genome.&lt;br /&gt;
&lt;br /&gt;
* Shallow-coverage (&amp;lt;10x) : you cannot use filtering to reliably separate TPs from FPs.  You must use the protocol involving [[Variant quality score recalibration]]&lt;br /&gt;
&lt;br /&gt;
The maximum DP (depth) filter only applies to whole genome data, where the probability of a site having exactly N reads given an average coverage of M is a well-behaved function.  First principles suggest this should be a binomial sampling but in practice it is more a Gaussian distribution.  Regardless, the DP threshold should be set a 5 or 6 sigma from the mean coverage across all samples, so that the DP &amp;gt; X threshold eliminates sites with excessive coverage caused by alignment artifacts.  Note that &#039;&#039;&#039;For exomes, a straight DP filter shouldn&#039;t be used&#039;&#039;&#039; because the relationship between misalignments and depth isn&#039;t clear for capture data.  Note that as of Dec. 2010 we no longer recommend using AB to filter, as (1) it&#039;s not generated by the UG by default any more and (2) QD is a better, and largely equivalent, filter.&lt;br /&gt;
&lt;br /&gt;
That said, all of the caveats about determining the right parameters, etc, are annoying and are largely eliminated by [[Variant quality score recalibration]].&lt;br /&gt;
&lt;br /&gt;
=== Making analysis ready calls with variant quality score recalibration ===&lt;br /&gt;
&lt;br /&gt;
An alternative approach to hard filtering, developed over the last 6 months in GSA to handle the low-pass whole genome data sets from the 1000 Genomes Project is an approach called [[Variant quality score recalibration]], described in detail on that wiki.  Guidelines for applying the [[Variant quality score recalibration]] are listed on the wiki, including which annotations to use and how to slice the recalibrated call set into the high-quality call sets.  The input call set to [[Variant quality score recalibration]] should be the snpFiltered.indelFiltered.vcf.  Note that the HARD_TO_VALIDATE filter are respected during training of the error model but these SNPs will be scored a potentially true variants, as this approach appears to robustly differentiate high-quality calls in the HARD_TO_VALIDATE set from the unreliable ones.&lt;br /&gt;
&lt;br /&gt;
== Expected SNP call quality ==&lt;br /&gt;
&lt;br /&gt;
All of the following tables were generated using [[VariantEval]].  You can generate your own data points for comparing with that tool and the correct input / comparison VCF files.&lt;br /&gt;
&lt;br /&gt;
=== Summary results for deep whole genome, multi-sample low-pass, and whole exome ===&lt;br /&gt;
&lt;br /&gt;
[[File:best_practice_calls.jpg|400px|thumb|Summary results for best practice calling pipeline]]&lt;br /&gt;
&lt;br /&gt;
The associated table provides some expectations for running deep whole genomes, single whole exomes, and multi-sample low-pass (from the 1000 genomes) in terms of sensitivity, specificity, and Ti/Tv ratios for known and novel calls.  Obviously individual data sets will be different but this demonstrates that running the pipeline described here ( realigned -&amp;gt; recalibration -&amp;gt; calling -&amp;gt; filtering -&amp;gt; variant quality score recalibration) can produce excellent results for a variety of data types.  Note that the deep data sets are single sample; in our hands multi-sample deep data results in much better calls overall.&lt;br /&gt;
&lt;br /&gt;
=== Expected Ti/Tv ratios ===&lt;br /&gt;
&lt;br /&gt;
[[File:titv_expectations.jpg|400px|thumb|Expected known and novel Ti/Tv ratios]]&lt;br /&gt;
&lt;br /&gt;
We provide two useful data points for evaluating the quality of SNP calls whole genome or in the targeted whole exome (Agilent).  You can follow a similar methodology to establish the Ti/Tv expectations for your region of interest, if captured separately, by running [[VariantEval]] on the 1000 Genomes Trios or Complete Genomes or other highly reliable data sets given the targeted interval.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Previous versions of Best Practices (now outdated) ==&lt;br /&gt;
===Version 1 -- hard filters===&lt;br /&gt;
* [[Whole exome v1]]&lt;br /&gt;
* [[Whole genome, deep coverage v1]]&lt;br /&gt;
* [[Whole genome, low-pass v1]]&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Sandbox:Best_Practice_Variant_Detection_with_the_GATK_v2&amp;diff=95005</id>
		<title>Sandbox:Best Practice Variant Detection with the GATK v2</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Sandbox:Best_Practice_Variant_Detection_with_the_GATK_v2&amp;diff=95005"/>
		<updated>2011-05-18T16:40:01Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: Created page with &amp;quot;== Data Processing Pipeline Script == The Data Processing Pipeline is a Queue script that performs all data processing described in this page following our most current recom...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Data Processing Pipeline Script ==&lt;br /&gt;
The [[Data Processing Pipeline]] is a Queue script that performs all data processing described in this page following our most current recommendations for best practices.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
Our current best practice for making SNP calls is divided into 5 sequential steps (see Figure): initial mapping, refinement of the initial reads, multi-sample indel and SNP calling, filtering of the raw SNP calls, and finally variant quality score recalibration.  These steps are the same for targeted resequencing, whole exomes, deep whole genomes, and low-pass whole genomes.  The exact commands for each tool are available on the tools wiki entry.   &lt;br /&gt;
&lt;br /&gt;
Note that, due to the specific attributes of a project, that the the specific values used in each of the commands may need to be selected by the analyst.  Care should be taken by the analyst running our tools to understand what each parameter does and to evaluate which value best fits his/her data.&lt;br /&gt;
&lt;br /&gt;
[[File:overall_flow.jpg|300px|thumb|Overview of GATK data processing pipeline]]&lt;br /&gt;
&lt;br /&gt;
=== Lanes, Samples, Cohort ===&lt;br /&gt;
&lt;br /&gt;
There are four major organizational units for next-generation DNA sequencing processes:&lt;br /&gt;
&lt;br /&gt;
; Lane : The basic machine unit for sequencing.  The lane reflects the basic independent run of an NGS machine.  For Illumina machines, this is the physical sequencing lane.  &lt;br /&gt;
&lt;br /&gt;
; Library : A unit of DNA preparation that at some point is physically pooled together.  Multiple lanes can be run from aliquots from the same library.  The DNA library and its preparation is the natural unit that is being sequenced.  For example, if the library has limited complexity, then many sequences are duplicated and will result in a high duplication rate across lanes. &lt;br /&gt;
&lt;br /&gt;
; Sample : A single individual, such as human CEPH NA12878.  Multiple libraries with different properties can be constructed from the original sample DNA source.  Here we treat samples as independent individuals whose genome sequence we are attempting to determine.  From this perspective, tumor / normal samples are different despite coming from the same individual.&lt;br /&gt;
&lt;br /&gt;
; Cohort : A collection of samples being analyzed together.  This organizational unit is the most subjective and depends intimately on the design goals of the sequencing project.  For population discovery projects like the 1000 Genomes, the analysis cohort is the ~100 individual in each population.  For exome projects with many samples (e.g., ESP with 800 EOMI samples) deeply sequenced we divide up the complete set of samples into cohorts of ~50 individuals for multi-sample analyses.  &lt;br /&gt;
&lt;br /&gt;
This document describes how to call variation within a single analysis cohort, comprised for one or many samples, each of one or many libraries that were sequenced on at least one lane of an NGS machine. &lt;br /&gt;
&lt;br /&gt;
Note that many GATK commands can be run at the lane level, but will give better results seeing all of the data for a single sample, or even all of the data for all samples.  Unfortunately, there&#039;s a trade-off in computational cost by running these commands across all of your data simultaneously.&lt;br /&gt;
&lt;br /&gt;
=== Testing data: 64x HiSeq on chr20 for NA12878 ===&lt;br /&gt;
&lt;br /&gt;
In order to help individuals get up to speed, evaluate their command lines, and generally become familiar with the GATK tools we recommend you download the raw and realigned, recalibrated [[NA12878 test data]].  It should be possible to apply all of the approaches outlined below to get excellent results for realignment, recalibration, SNP calling, indel calling, filtering, and variant quality score recalibration using this data.&lt;br /&gt;
&lt;br /&gt;
== Phase I: Raw data processing ==&lt;br /&gt;
[[Image:DataProcessingPPL.jpg|thumb|300px|Data processing pipeline of the best practices for raw data processing, from sequencer data (fastq files) to analysis read reads (bam file)]]&lt;br /&gt;
&lt;br /&gt;
=== Initial read mapping ===&lt;br /&gt;
&lt;br /&gt;
The GATK data processing pipeline assumes that one of the many NGS read aligners (see [http://bib.oxfordjournals.org/content/11/5/473.abstract] for a review) has been applied to your raw FASTQ files.  For Illumina data we recommend [http://bio-bwa.sourceforge.net/ BWA] because it is accurate, fast, well-supported, open-source, and emits BAM files natively.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Raw BAM to realigned, recalibrated BAM ===&lt;br /&gt;
&lt;br /&gt;
The three key tools here are [[Base quality score recalibration]], [[Local realignment around indels]], and [http://picard.sourceforge.net/command-line-overview.shtml#MarkDuplicates|Picard&#039;s MarkDuplicates].  Although ideally one follows the flow from Figure 1, in practice MarkDuplicates can be run before local realignment, in order to handle cases where duplicates overlapping indels get marginally different alignments   (unlikely but possible) and so will not be considered as potential duplicates because MarkDuplicates looks only at read pair start and stop positions.  There are several options here, from the easy and fast basic protocol to the more.&lt;br /&gt;
&lt;br /&gt;
There are two types of realignment:&lt;br /&gt;
* Realignment only at known sites, which is very efficient, can operate with little coverage (1x per lane genome wide) but can only realign reads at known indels.&lt;br /&gt;
* Fully local realignment uses mismatching bases to determine if a site should be realigned, and relies on sufficient coverage to discover the correct indel allele in the reads for alignment.  It is much slower (involves SW step) but can discover new indel sites in the reads.  If you have a database of known indels (for human, this database is extensive) then at this stage you would also include these indels during realignment, which vastly improves sensitivity, specificity, and speed.&lt;br /&gt;
&lt;br /&gt;
==== Previous recommendation: lane-level recalibration, sample-level realignment ====&lt;br /&gt;
&lt;br /&gt;
This is the protocol we used at the Broad Institute for the last year.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each lane.bam&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicate(lane.bam)&lt;br /&gt;
    recal.bam &amp;lt;- recal(dedup.bam)&lt;br /&gt;
&lt;br /&gt;
for each sample&lt;br /&gt;
    lanes.bam &amp;lt;- merged recal.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(lanes.bam)&lt;br /&gt;
    realigned.bam &amp;lt;- realign(dedup.bam) [with known sites if possible]&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fast: lane-level realignment at known sites only and lane-level recalibration ====&lt;br /&gt;
&lt;br /&gt;
This protocol adds lane-level local realignment around known indels, making it very fast (no sample level processing) and gives better results for human samples than the previous recommendation:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each lane.bam&lt;br /&gt;
    realigned.bam &amp;lt;- realign(lane.bam) [at only known sites]&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicate(realigned.bam)&lt;br /&gt;
    recal.bam &amp;lt;- recal(dedup.bam)&lt;br /&gt;
&lt;br /&gt;
for each sample&lt;br /&gt;
    recals.bam &amp;lt;- merged lane-level recal.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(recals.bam)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fast + sample-level realignment ====&lt;br /&gt;
&lt;br /&gt;
This protocol adds sample-level realignment after library / sample level dedupping, so that novel indels in each sample can be discovered and realigned around.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each lane.bam&lt;br /&gt;
    realigned.bam &amp;lt;- realign(lane.bam) [at only known sites]&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicate(realigned.bam)&lt;br /&gt;
    recal.bam &amp;lt;- recal(dedup.bam)&lt;br /&gt;
&lt;br /&gt;
for each sample&lt;br /&gt;
    recals.bam &amp;lt;- merged lane-level recal.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(recals.bam)&lt;br /&gt;
    realigned.bam &amp;lt;- realign(dedup.bam) [with known sites if possible]&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Better: sample-level realignment with known indels and recalibration ====&lt;br /&gt;
&lt;br /&gt;
Rather than doing the lane level cleaning and recalibration, this process aggregates all of the reads for each sample and then does a full dedupping, realign, and recalibration, yielding the best single-sample results.  The big change here is sample-level cleaning followed by recalibration, giving you the most accurate quality scores possible for a single sample. &lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each sample&lt;br /&gt;
    lanes.bam &amp;lt;- merged lane.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(lanes.bam)&lt;br /&gt;
    realigned.bam &amp;lt;- realign(dedup.bam) [with known sites if possible]&lt;br /&gt;
    recal.bam &amp;lt;- recal(realigned.bam)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Best: multi-sample realignment with known sites and recalibration ====&lt;br /&gt;
&lt;br /&gt;
Finally, if you really want to get the absolute best results, whatever the computational cost, then we recommend doing multiple sample realignment so that novel indels in one sample help to realign reads in other samples.  Although not generally necessary for deep sequencing data, this is important for low-coverage multi-sample SNP calling projects like the 1000 Genomes Project.  Note that the computational cost here is so extreme that we only do this analysis in special circumstances, such as large-scale data freeze for the project.  &lt;br /&gt;
&lt;br /&gt;
Note that for contrastive calling projects -- such as cancer tumor/normals -- that we recommend cleaning both the tumor and the normal together in general to avoid slight alignment differences between the two tissue types.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
for each sample&lt;br /&gt;
    lanes.bam &amp;lt;- merged lane.bam&#039;s for sample&lt;br /&gt;
    dedup.bam &amp;lt;- MarkDuplicates(lanes.bam)&lt;br /&gt;
&lt;br /&gt;
samples.bam &amp;lt;- merged dedup.bam&#039;s for all samples&lt;br /&gt;
realigned.bam &amp;lt;- realign(samples.bam)&lt;br /&gt;
recal.bam &amp;lt;- recal(realigned.bam)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Misc. notes on the process ====&lt;br /&gt;
&lt;br /&gt;
* MarkDuplicates needs only be run at the library level. So the sample-level dedupping isn&#039;t necessary if you only ever a library on a single lane.  If you run the sample library on many lanes (as can be necessary for whole exome, for example), you should dedup at the library level.&lt;br /&gt;
* The base quality score recalibrator is read group aware, so running it on a merged BAM files containing multiple read groups is the same as running it on each bam file individually.  There&#039;s some memory cost (so it&#039;s best not to recalibrate 10000 RGs simultaneously) but for reasonable projects this is a fine.&lt;br /&gt;
* Local realignment preserves read meta-data, so you can realign and then recalibrate just fine.&lt;br /&gt;
&lt;br /&gt;
== Initial variant discovery and genotyping ==&lt;br /&gt;
&lt;br /&gt;
=== Input BAMs for variant discovery and genotyping ===&lt;br /&gt;
&lt;br /&gt;
After the raw data processing step, the GATK variant detection process assumes that you have aligned, duplicate marked, and recalibrated BAM files for all of the samples in your cohort.  Because the GATK can dynamically merge BAM files, it isn&#039;t critical to have merged files by lane into sample bams, or even samples bams into cohort bams.  In general we try to create sample level bams for deep data sets (deep WG or exomes) and merged cohort files by chromosome for WG low-pass.  A nice size for BAMs is 10-300 Gb, just for organizing on disk.  &lt;br /&gt;
&lt;br /&gt;
For this part of the this document, I&#039;m going to assume that you have a single realigned, recalibrated, dedupped BAM per sample, called sampleX.bam, for X from 1 to N samples in your cohort.  Note that some of the data processing steps, such as multiple sample local realignment, will merge BAMS for many samples into a single BAM.  If you&#039;ve gone down this route, you just need to modify the GATK commands as necessary to take not multiple BAMs, one for each sample, but a single BAM for all samples.&lt;br /&gt;
&lt;br /&gt;
=== Multi-sample SNP and indel calling ===&lt;br /&gt;
&lt;br /&gt;
The next step in the standard GATK data processing pipeline, whole genome or targeted, deep or shallow, is to apply the [[Unified genotyper]] to identify sites among the cohort samples that are statistically non-reference.  This will produce a multi-sample [[VCF Format|VCF file]], with sites discovered across samples and genotypes assigned to each sample in the cohort.  It&#039;s in this stage that we use the meta-data in the BAM files extensively -- read groups for reads, with samples, platforms, etc -- to enable us to do the multi-sample merging and genotyping correctly.  It was a pain for data processing, yes, but now life is easy for downstream calling and analysis.&lt;br /&gt;
&lt;br /&gt;
Note that by default the Unified Genotyper calls SNPs only.  To enable the indel calling capabilities instead use the -glm DINDEL argument.&lt;br /&gt;
&lt;br /&gt;
==== Selecting an appropriate quality score threshold ====&lt;br /&gt;
&lt;br /&gt;
A common question is the confidence score threshold to use for variant detection.  We recommend:&lt;br /&gt;
&lt;br /&gt;
; Deep (&amp;gt; 10x coverage per sample) data : we recommend a minimum confidence score threshold of Q30 with an emission threshold of Q10.  These Q10-Q30 calls will be emitted filtered out as LowQual.  &lt;br /&gt;
; Shallow (&amp;lt; 10x coverage per sample) data : because variants have by necessity lower quality with shallower coverage, we recommend a min. confidence score of Q4 and an emission threshold of Q3. &lt;br /&gt;
&lt;br /&gt;
If you want include these LowQual variants following [[Variant quality score recalibration]] you can just ignore the LowQual filter in the VariantRecalibration step.  That way they won&#039;t influence the recalibration but can be selected, if they appear sufficiently good, for inclusion in the final call set.&lt;br /&gt;
&lt;br /&gt;
=== Protocol ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
raw.vcf &amp;lt;- unifiedGenotyper(sample1.bam, sample2.bam, ..., sampleN.bam)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Integrating analyses: getting the best call set possible ==&lt;br /&gt;
&lt;br /&gt;
This raw VCF file should be as sensitive to variation as you&#039;ll get without imputation.  At this stage, you can assess things like sensitivity to known variant sites or genotype chip concordance.  The problem is that the raw VCF will have many sites that aren&#039;t really genetic variants but are machine artifacts that make the site statistically non-reference.  All of the subsequent filtering and/or quality score recalibration steps are designed to separate out the FP machine artifacts from the TP genetic variants.&lt;br /&gt;
&lt;br /&gt;
The tools used here is [[VariantFiltrationWalker]] to apply hard filters and [[Variant quality score recalibration]] to build an adaptive error model using known variant sites and then apply this model to estimate the probability that each variant is a true genetic variant or a machine artifact.  Regardless of whether you&#039;ll ultimately apply hard filtering or adaptive error modeling to select your final calls, we recommend you first apply some common SNP filters to avoid obvious misalignment and indel artifacts.  &lt;br /&gt;
&lt;br /&gt;
=== Analysis read VCF procotol ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
snpFiltered.vcf &amp;lt;- basicSNPFilters(raw.vcf)&lt;br /&gt;
snpFiltered.indelFiltered.vcf &amp;lt;- indelFilters(snpFiltered.vcf)&lt;br /&gt;
&lt;br /&gt;
if we are hard filtering:&lt;br /&gt;
    analysisReady.vcf &amp;lt;- dataTypeSpecificFilters(snpFiltered.indelFiltered.vcf)&lt;br /&gt;
&lt;br /&gt;
else we are doing variant quality score recalibration&lt;br /&gt;
    analysisReady.vcf &amp;lt;- variantQualityScoreRecalibrate(snpFiltered.indelFiltered.vcf)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note that [[VariantFiltrationWalker]] can likely apply simultaneously all of the filters you want.  For easy of explanation, we divide up the application of these filters into independent applications of the [[VariantFiltrationWalker]]. Overall, the approximate command for variant filtering will look something like, but depending on exactly what how you want to filter your variants the detail of this command will change:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Common VariantFiltrationWalker command&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
java -jar GenomeAnalysisTK.jar \&lt;br /&gt;
  -T VariantFiltration \&lt;br /&gt;
  -R resources/Homo_sapiens_assembly18.fasta \&lt;br /&gt;
  -B:variant,VCF input.vcf \&lt;br /&gt;
  -o filtered.vcf \&lt;br /&gt;
  [FILTERING PARAMETERS]&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Basic indel filtering ===&lt;br /&gt;
&lt;br /&gt;
The goal of the Basic indel filters is to remove alignment artifacts from the data.  The key item here is flagging variants with high strand bias and in poorly mapped regions (HARD_TO_VALIDATE set) with more than 10% of the reads having mapping quality 0.&lt;br /&gt;
Arguments for [[VariantFiltrationWalker]]:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  --filterExpression &amp;quot;MQ0 &amp;gt;= 4 &amp;amp;&amp;amp; ((MQ0 / (1.0 * DP)) &amp;gt; 0.1)&amp;quot;    [expression to match 10% of reads with MAPQ0]&lt;br /&gt;
  --filterName &amp;quot;HARD_TO_VALIDATE&amp;quot;&lt;br /&gt;
  --filterExpression &amp;quot;SB &amp;gt;= -1.0&amp;quot;&lt;br /&gt;
  --filterName &amp;quot;StrandBiasFilter&amp;quot;&lt;br /&gt;
  --filterExpression &amp;quot;QUAL &amp;lt; 10&amp;quot;&lt;br /&gt;
  --filterName &amp;quot;QualFilter&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note that, for the above to work, the input vcf needs to be annotated with the corresponding values (MQ0,DP,SB). If any of these values are missing, then VariantAnnotator needs to be run first so that VariantFiltration can run properly.&lt;br /&gt;
&lt;br /&gt;
=== Basic SNP filtering ===&lt;br /&gt;
&lt;br /&gt;
The goal of the Basic SNP filters is to remove alignment artifacts from the data.  The key items here is flagging SNPs within clusters (3 SNPs with 10 bp of each other) and those in poorly mapped regions (HARD_TO_VALIDATE set) with more than 10% of the reads having mapping quality 0.&lt;br /&gt;
&lt;br /&gt;
Arguments for [[VariantFiltrationWalker]]:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  --clusterWindowSize 10 &lt;br /&gt;
  --filterExpression &amp;quot;MQ0 &amp;gt;= 4 &amp;amp;&amp;amp; ((MQ0 / (1.0 * DP)) &amp;gt; 0.1)&amp;quot;    [expression to match 10% of reads with MAPQ0]&lt;br /&gt;
  --filterName &amp;quot;HARD_TO_VALIDATE&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The clusterWindowSize can be reduced below 10 if you expect many true variants close together, such as in an organisms with high \math{theta} or with many individuals (1000s of humans).&lt;br /&gt;
&lt;br /&gt;
As mentioned above, this command assumes that MQ0 and DP are present in the input vcf. If they are not, then you need to run VariantAnnotator first to complete the corresponding annotations.&lt;br /&gt;
&lt;br /&gt;
=== Filtering around indels ===&lt;br /&gt;
&lt;br /&gt;
It&#039;s possible that, despite even local realignment, misalignments around true and artifactual indels will result in some false SNP calls.  These errors are quite common if you didn&#039;t do local realignment, didn&#039;t provide a set of known indels during local realignment, and around very large indels that can&#039;t be modeled properly by local realignment.  If you performed indel calling then you can filter your SNP calls around the raw indel calls from your data set.  The raw calls should be used so that you can remove SNPs around even artifactual indels, which even though the indel caller doesn&#039;t believe they are real, are sufficient to generate FP SNP calls.  Skipping this step will leave some variants in your call set that would have been identified as indel artifacts.  It&#039;s possible that downstream filters or variant quality score recalibration will remove these, but if you want the best calls possible, it&#039;s worth the effort to indel call for the mask alone.&lt;br /&gt;
&lt;br /&gt;
Arguments for [[VariantFiltrationWalker]]:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  -B:mask,VCF /path/to/indels.from.UnifiedGenotyper.vcf&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
If you have good indel calls for your population (for example, the 1000 Genomes Project pilot indel calls that include 1.3M sites) you can filter calls around these indels by generating a mask for those indels.  This is a more advanced approach that&#039;s only really necessary if you don&#039;t realign around indels, but can help avoid even more indel artifacts.  If these indels are not available in VCF format, you can create a [http://genome.ucsc.edu/FAQ/FAQformat.html#format1 Bed] file with the loci to mask, and supply the mask with the following command:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  -B:mask,Bed /path/to/indels.mask.bed&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Recall that the bed format is 0-based and half-closed.&lt;br /&gt;
&lt;br /&gt;
; Deep (&amp;gt; 10x coverage per sample) data : we recommend the raw indel calls from UnifiedGenotyper be supplied as the mask.&lt;br /&gt;
; Shallow (&amp;lt; 10x coverage per sample) data : given the soft-calling procedure, which can learn the appropriate filters automatically, an indel mask is not required.&lt;br /&gt;
&lt;br /&gt;
=== Making analysis ready calls SNP calls with hard filtering ===&lt;br /&gt;
&lt;br /&gt;
We include these hard filter parameters for several reasons.  One, they give reasonably good results with sufficient depth and multiple samples.  Second, many data sets were generated and filtered with these parameters and we want to make sure that people can generate similar data sets in the future.  Finally, VQSR requires lots of data to work properly, and with targeted resequencing of a small region (for example, a few hundred genes) VQSR may be inappropriate leaving hard filtering as the only option.  &lt;br /&gt;
&lt;br /&gt;
Arguments for [[VariantFiltrationWalker]]:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  --filterExpression [DATA_TYPE_SPECIFIC_FILTERS]&lt;br /&gt;
  --filterName GATKStandard&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where DATA_TYPE_SPECIFIC_FILTERS has project-specific filtering strings selected below:&lt;br /&gt;
&lt;br /&gt;
* For exomes with deep coverage per sample &lt;br /&gt;
** DATA_TYPE_SPECIFIC_FILTERS should be &amp;quot;QUAL &amp;lt; 30.0 || QD &amp;lt; 5.0 || HRun &amp;gt; 5 || SB &amp;gt; -0.10&amp;quot;&lt;br /&gt;
&lt;br /&gt;
* Whole genomes with deep coverage:  &lt;br /&gt;
** DATA_TYPE_SPECIFIC_FILTERS should be &amp;quot;DP &amp;gt; 100 || MQ0 &amp;gt; 40 || SB &amp;gt; -0.10&amp;quot;. &lt;br /&gt;
** Note that the DP values (depth filters) must be parameterized by the average coverage among all of the samples.  That is, if you have 100x mean coverage, there&#039;s nothing wrong with a site with 110x, but usually 1000x indicates a bad sites, at least whole genome.&lt;br /&gt;
&lt;br /&gt;
* Shallow-coverage (&amp;lt;10x) : you cannot use filtering to reliably separate TPs from FPs.  You must use the protocol involving [[Variant quality score recalibration]]&lt;br /&gt;
&lt;br /&gt;
The maximum DP (depth) filter only applies to whole genome data, where the probability of a site having exactly N reads given an average coverage of M is a well-behaved function.  First principles suggest this should be a binomial sampling but in practice it is more a Gaussian distribution.  Regardless, the DP threshold should be set a 5 or 6 sigma from the mean coverage across all samples, so that the DP &amp;gt; X threshold eliminates sites with excessive coverage caused by alignment artifacts.  Note that &#039;&#039;&#039;For exomes, a straight DP filter shouldn&#039;t be used&#039;&#039;&#039; because the relationship between misalignments and depth isn&#039;t clear for capture data.  Note that as of Dec. 2010 we no longer recommend using AB to filter, as (1) it&#039;s not generated by the UG by default any more and (2) QD is a better, and largely equivalent, filter.&lt;br /&gt;
&lt;br /&gt;
That said, all of the caveats about determining the right parameters, etc, are annoying and are largely eliminated by [[Variant quality score recalibration]].&lt;br /&gt;
&lt;br /&gt;
=== Making analysis ready calls with variant quality score recalibration ===&lt;br /&gt;
&lt;br /&gt;
An alternative approach to hard filtering, developed over the last 6 months in GSA to handle the low-pass whole genome data sets from the 1000 Genomes Project is an approach called [[Variant quality score recalibration]], described in detail on that wiki.  Guidelines for applying the [[Variant quality score recalibration]] are listed on the wiki, including which annotations to use and how to slice the recalibrated call set into the high-quality call sets.  The input call set to [[Variant quality score recalibration]] should be the snpFiltered.indelFiltered.vcf.  Note that the HARD_TO_VALIDATE filter are respected during training of the error model but these SNPs will be scored a potentially true variants, as this approach appears to robustly differentiate high-quality calls in the HARD_TO_VALIDATE set from the unreliable ones.&lt;br /&gt;
&lt;br /&gt;
== Expected SNP call quality ==&lt;br /&gt;
&lt;br /&gt;
All of the following tables were generated using [[VariantEval]].  You can generate your own data points for comparing with that tool and the correct input / comparison VCF files.&lt;br /&gt;
&lt;br /&gt;
=== Summary results for deep whole genome, multi-sample low-pass, and whole exome ===&lt;br /&gt;
&lt;br /&gt;
[[File:best_practice_calls.jpg|400px|thumb|Summary results for best practice calling pipeline]]&lt;br /&gt;
&lt;br /&gt;
The associated table provides some expectations for running deep whole genomes, single whole exomes, and multi-sample low-pass (from the 1000 genomes) in terms of sensitivity, specificity, and Ti/Tv ratios for known and novel calls.  Obviously individual data sets will be different but this demonstrates that running the pipeline described here ( realigned -&amp;gt; recalibration -&amp;gt; calling -&amp;gt; filtering -&amp;gt; variant quality score recalibration) can produce excellent results for a variety of data types.  Note that the deep data sets are single sample; in our hands multi-sample deep data results in much better calls overall.&lt;br /&gt;
&lt;br /&gt;
=== Expected Ti/Tv ratios ===&lt;br /&gt;
&lt;br /&gt;
[[File:titv_expectations.jpg|400px|thumb|Expected known and novel Ti/Tv ratios]]&lt;br /&gt;
&lt;br /&gt;
We provide two useful data points for evaluating the quality of SNP calls whole genome or in the targeted whole exome (Agilent).  You can follow a similar methodology to establish the Ti/Tv expectations for your region of interest, if captured separately, by running [[VariantEval]] on the 1000 Genomes Trios or Complete Genomes or other highly reliable data sets given the targeted interval.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Previous versions of Best Practices (now outdated) ==&lt;br /&gt;
===Version 1 -- hard filters===&lt;br /&gt;
* [[Whole exome v1]]&lt;br /&gt;
* [[Whole genome, deep coverage v1]]&lt;br /&gt;
* [[Whole genome, low-pass v1]]&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Integrated_Sciences_Student_Association&amp;diff=94955</id>
		<title>Integrated Sciences Student Association</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Integrated_Sciences_Student_Association&amp;diff=94955"/>
		<updated>2011-05-18T00:42:39Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: Created the page.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Who We Are==&lt;br /&gt;
&lt;br /&gt;
The Integrated Sciences Student Association (ISSA) provides a means for social interaction amongst Integrated Science (ISCI) students.  This is especially important for ISCI because of the nature of the program.  The diversity of foci explored by ISCI students means that they do not necessarily share many classes with fellow ISCI students.  ISSA is here to bridge the gap between students the same way ISCI students bridge the gap between sciences.&lt;br /&gt;
&lt;br /&gt;
ISSA also promotes the Integrated Sciences Program, spreading awareness about the wealth options available to students who choose to pursue multidisciplinary studies in the sciences.&lt;br /&gt;
&lt;br /&gt;
==Our Website==&lt;br /&gt;
&lt;br /&gt;
For more information about ISSA, including our upcoming events and current executives, visit http://www.ubcissa.com.&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=User:Ezhao&amp;diff=94954</id>
		<title>User:Ezhao</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=User:Ezhao&amp;diff=94954"/>
		<updated>2011-05-18T00:37:08Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Welcome to my page==&lt;br /&gt;
&lt;br /&gt;
My name is Eric Zhao, and I am a third year student in Physiology with a proposed minor in Physics.&lt;br /&gt;
&lt;br /&gt;
I am excited to be a part of UBC wiki and hope that it will be a wonderful shared resource for all.  I hope that teachers get actively involved, as I see this as an opportunity to write out unhelpful and expensive textbooks in favour of streamlined, relevant, and open-source learning resources.&lt;br /&gt;
&lt;br /&gt;
Here are some pages I&#039;m involved with:&lt;br /&gt;
&lt;br /&gt;
* [[Safety and First Aid Club]]&lt;br /&gt;
* [[Integrated Sciences Student Association]]&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=User:Ezhao&amp;diff=94953</id>
		<title>User:Ezhao</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=User:Ezhao&amp;diff=94953"/>
		<updated>2011-05-18T00:36:45Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Welcome to my page==&lt;br /&gt;
&lt;br /&gt;
My name is Eric Zhao, and I am a third year student in Physiology with a proposed minor in Physics.&lt;br /&gt;
&lt;br /&gt;
I am excited to be a part of UBC wiki and hope that it will be a wonderful shared resource for all.  I hope that teachers get actively involved, as I see this as an opportunity to write out unhelpful and expensive textbooks in favour of streamlined, relevant, and open-source learning resources.&lt;br /&gt;
&lt;br /&gt;
Here are some pages I&#039;m involved with:&lt;br /&gt;
&lt;br /&gt;
* &#039;Safety and First Aid Club&#039;&lt;br /&gt;
* &#039;Integrated Sciences Student Association&#039;&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=User:Ezhao&amp;diff=94952</id>
		<title>User:Ezhao</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=User:Ezhao&amp;diff=94952"/>
		<updated>2011-05-18T00:36:32Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Welcome to my page==&lt;br /&gt;
&lt;br /&gt;
My name is Eric Zhao, and I am a third year student in Physiology with a proposed minor in Physics.&lt;br /&gt;
&lt;br /&gt;
I am excited to be a part of UBC wiki and hope that it will be a wonderful shared resource for all.  I hope that teachers get actively involved, as I see this as an opportunity to write out unhelpful and expensive textbooks in favour of streamlined, relevant, and open-source learning resources.&lt;br /&gt;
&lt;br /&gt;
Here are some pages I&#039;m involved with:&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;Safety and First Aid Club&#039;&#039;&lt;br /&gt;
* &#039;&#039;Integrated Sciences Student Association&#039;&#039;&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Safety_and_First_Aid_Club&amp;diff=94951</id>
		<title>Safety and First Aid Club</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Safety_and_First_Aid_Club&amp;diff=94951"/>
		<updated>2011-05-18T00:34:59Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Who We Are==&lt;br /&gt;
The UBC Safety and First Aid Club is dedicated to bringing awareness and skills in first aid to all people on campus and in the community.  A small group of students, we do this by offering affordable first aid courses on campus and planning outreach programs to educate students of Vancouver secondary schools.&lt;br /&gt;
&lt;br /&gt;
==Our Courses==&lt;br /&gt;
&lt;br /&gt;
We offer certification in Standard First Aid with CPR-C with our courses.  Each course runs from 9-7 on a Saturday, and passing students receive their certification upon completion.  We typically hold one course each month.  Each course costs $90 ($5 membership included).&lt;br /&gt;
&lt;br /&gt;
====No-fail Policy====&lt;br /&gt;
&lt;br /&gt;
We implement a no-fail policy.  Any student who does not pass a course may attend the next one for free.&lt;br /&gt;
&lt;br /&gt;
====How to Sign Up====&lt;br /&gt;
&lt;br /&gt;
For course information, please visit our website at http://www.sfasociety.org.  To sign up, visit our office or get in contact with an executive (see website for contact information).  We require that all students pay a $20 deposit to ensure their spot in the course.  The remaining $70 is collected on the course day.&lt;br /&gt;
&lt;br /&gt;
==Getting Involved==&lt;br /&gt;
&lt;br /&gt;
Aside from president and vice-president positions, all other executive positions are filled by application.  We always welcome positive contributions by enthusiastic members.  Get in contact in order to apply.&lt;br /&gt;
&lt;br /&gt;
==Our Website==&lt;br /&gt;
&lt;br /&gt;
Visit http://www.sfasociety.org for more information.&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Safety_and_First_Aid_Club&amp;diff=94950</id>
		<title>Safety and First Aid Club</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Safety_and_First_Aid_Club&amp;diff=94950"/>
		<updated>2011-05-18T00:34:42Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: Created the page.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Who We Are==&lt;br /&gt;
The UBC Safety and First Aid Club is dedicated to bringing awareness and skills in first aid to all people on campus and in the community.  A small group of students, we do this by offering affordable first aid courses on campus and planning outreach programs to educate students of Vancouver secondary schools.&lt;br /&gt;
&lt;br /&gt;
==Our Courses==&lt;br /&gt;
&lt;br /&gt;
We offer certification in Standard First Aid with CPR-C with our courses.  Each course runs from 9-7 on a Saturday, and passing students receive their certification upon completion.  We typically hold one course each month.  Each course costs $90 ($5 membership included).&lt;br /&gt;
&lt;br /&gt;
===No-fail Policy===&lt;br /&gt;
&lt;br /&gt;
We implement a no-fail policy.  Any student who does not pass a course may attend the next one for free.&lt;br /&gt;
&lt;br /&gt;
===How to Sign Up===&lt;br /&gt;
&lt;br /&gt;
For course information, please visit our website at http://www.sfasociety.org.  To sign up, visit our office or get in contact with an executive (see website for contact information).  We require that all students pay a $20 deposit to ensure their spot in the course.  The remaining $70 is collected on the course day.&lt;br /&gt;
&lt;br /&gt;
==Getting Involved==&lt;br /&gt;
&lt;br /&gt;
Aside from president and vice-president positions, all other executive positions are filled by application.  We always welcome positive contributions by enthusiastic members.  Get in contact in order to apply.&lt;br /&gt;
&lt;br /&gt;
==Our Website==&lt;br /&gt;
&lt;br /&gt;
Visit http://www.sfasociety.org for more information.&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=List_of_Vancouver_Clubs&amp;diff=94949</id>
		<title>List of Vancouver Clubs</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=List_of_Vancouver_Clubs&amp;diff=94949"/>
		<updated>2011-05-18T00:23:47Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: Added safety and first aid club.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==AMS Affiliated Clubs==&lt;br /&gt;
&lt;br /&gt;
=== A ===&lt;br /&gt;
* [[A Cappella Club]]&lt;br /&gt;
* [[Accounting Club]]&lt;br /&gt;
* [[Africa Awareness Initiative]]&lt;br /&gt;
* [[Agents for Change]]&lt;br /&gt;
* [[AIDS Community Action Network]]&lt;br /&gt;
* [[Alternative &amp;amp; Integrative Medicine Society]]&lt;br /&gt;
* [[Alternative Arts]]&lt;br /&gt;
* [[Amateur Radio Society of UBC]]&lt;br /&gt;
* [[Ambassadors for Jesus]]&lt;br /&gt;
* [[American Sign Language Club of UBC]]&lt;br /&gt;
* [[Amnesty Iternational UBC]]&lt;br /&gt;
* [[Anime Club]]&lt;br /&gt;
* [[Anthropology Students Association]]&lt;br /&gt;
* [[Aqua Society]]&lt;br /&gt;
* [[Arab Students&#039; Association]]&lt;br /&gt;
* [[Art History Students&#039; Association]]&lt;br /&gt;
* [[Arts Co-op Students&#039; Association]]&lt;br /&gt;
* [[Asian-Canadian Cultural Organization]]&lt;br /&gt;
* [[Asian Debate Club]]&lt;br /&gt;
* [[Asian Studies Interests Association]]&lt;br /&gt;
* [[Assembly of Interdisciplinary Students]]&lt;br /&gt;
* [[Association International des Etudiants en Sciences Economiques et Commerciales]]&lt;br /&gt;
* [[Association of Canadian Archivists]]&lt;br /&gt;
* [[Association of Chinese Graduates]]&lt;br /&gt;
* [[Association of Korean-Canadian Scientists &amp;amp; Engineers]]&lt;br /&gt;
* [[Association of Latin American Students]]&lt;br /&gt;
* [[Astronomy Club of UBC]]&lt;br /&gt;
* [[Aviation Club]]&lt;br /&gt;
&lt;br /&gt;
=== B ===&lt;br /&gt;
* [[Badminton Club of the AMS]]&lt;br /&gt;
* [[Bangladesh Students&#039; Association]]&lt;br /&gt;
* [[BC Young Liberals]]&lt;br /&gt;
* [[Beads &amp;amp; Crafts Club]]&lt;br /&gt;
* [[Best Buddies UBC]]&lt;br /&gt;
* [[Bhangra Club]]&lt;br /&gt;
* [[Bike Co-Op]]&lt;br /&gt;
* [[Biochemistyr Pharmacology Physiology Club]]&lt;br /&gt;
* [[Biological Sciences Society]]&lt;br /&gt;
* [[BioSynergy at UBC]]&lt;br /&gt;
* [[Born for More Christian Student Ministries]]&lt;br /&gt;
* [[Botany Enthusiasts Club]]&lt;br /&gt;
* [[Brazilian Connection]]&lt;br /&gt;
* [[Business Asia]]&lt;br /&gt;
* [[Business Communications]]&lt;br /&gt;
* [[Business Technology Network]]&lt;br /&gt;
&lt;br /&gt;
=== C ===&lt;br /&gt;
* [[Campus Association for Baha&#039;i Studies]]&lt;br /&gt;
* [[Campus for Christ]]&lt;br /&gt;
* [[Canadian Association of Pharmacy Students &amp;amp; Interns]]&lt;br /&gt;
* [[Canadian Students for Sensible Drug Policy]]&lt;br /&gt;
* [[Cancer Association]]&lt;br /&gt;
* [[Caribbean African Association]]&lt;br /&gt;
* [[Chess Club]]&lt;br /&gt;
* [[Chin Woo Athletic Association]]&lt;br /&gt;
* [[Chinese Art Students&#039; Society]]&lt;br /&gt;
* [[Chinese Catholic Society]]&lt;br /&gt;
* [[Chinese Chess Club]]&lt;br /&gt;
* [[Chinese Christian Fellowship]]&lt;br /&gt;
* [[Chinese Collegiate Society]]&lt;br /&gt;
* [[Chinese Students&#039; Association]]&lt;br /&gt;
* [[Chinese Varsity Club]]&lt;br /&gt;
* [[Chiropractic Club]]&lt;br /&gt;
* [[Choy Lee Fat Kung Fu]]&lt;br /&gt;
* [[Christian Students at UBC]]&lt;br /&gt;
* [[Circle K (Kiwanis) Volunteers]]&lt;br /&gt;
* [[Civil Engineering Club]]&lt;br /&gt;
* [[Classical, Near Eastern &amp;amp; Religious Studies Student Association]]&lt;br /&gt;
* [[Cognitive Systems Society]]&lt;br /&gt;
* [[Coin &amp;amp; Stamp Club]]&lt;br /&gt;
* [[Commonwealth Club]]&lt;br /&gt;
* [[Competitive Video Gaming Association]]&lt;br /&gt;
* [[Computer Co-Op]]&lt;br /&gt;
* [[Computer Science Students&#039; Society]]&lt;br /&gt;
* [[Consulting Club]]&lt;br /&gt;
* [[Cycling Club]]&lt;br /&gt;
&lt;br /&gt;
=== D ===&lt;br /&gt;
* [[Dance Club]]&lt;br /&gt;
* [[Dance Horizons]]&lt;br /&gt;
* [[Dance Team]]&lt;br /&gt;
* [[Debating Society]]&lt;br /&gt;
* [[Dollar Project UBC]]&lt;br /&gt;
* [[Dragon Boating Club]]&lt;br /&gt;
* [[Dragon Seed Connection]]&lt;br /&gt;
&lt;br /&gt;
=== E ===&lt;br /&gt;
* [[Economics Students&#039; Association]]&lt;br /&gt;
* [[Electrical &amp;amp; Computer Engineering Student Society]]&lt;br /&gt;
* [[Emerging Leaders of UBC]]&lt;br /&gt;
* [[Engineering Physics Society]]&lt;br /&gt;
* [[English Students&#039; Association]]&lt;br /&gt;
* [[Environmental Sciences Students&#039; Association]]&lt;br /&gt;
* [[Environmental Design Society]]&lt;br /&gt;
* [[Equestrian Club]]&lt;br /&gt;
* [[European Studies Students&#039; Association]]&lt;br /&gt;
&lt;br /&gt;
=== F ===&lt;br /&gt;
* [[Federal Young Liberals of UBC]]&lt;br /&gt;
* [[Fencing Club]]&lt;br /&gt;
* [[Film Society]]&lt;br /&gt;
* [[Finance Club]]&lt;br /&gt;
* [[Food Society]]&lt;br /&gt;
* [[Formation DanceSport UBC]]&lt;br /&gt;
* [[Forum on International Cooperation]]&lt;br /&gt;
* [[Freethinkers Club]]&lt;br /&gt;
* [[Friends of the Spartacus Youth Club]]&lt;br /&gt;
* [[Friends of the UBC Farm]]&lt;br /&gt;
* [[Friends of Wetlands]]&lt;br /&gt;
* [[Fun Run UBC Running Club]]&lt;br /&gt;
&lt;br /&gt;
=== G ===&lt;br /&gt;
* [[Gado-Gado Indonesian Students&#039; Association of UBC]]&lt;br /&gt;
* [[Game Developers&#039; Association]]&lt;br /&gt;
* [[Geography Students&#039; Association]]&lt;br /&gt;
* [[German Club]]&lt;br /&gt;
* [[Gilbert &amp;amp; Sullivan Society of UBC]]&lt;br /&gt;
* [[Go Club]]&lt;br /&gt;
* [[Golf Club]]&lt;br /&gt;
* [[Green Party of UBC]]&lt;br /&gt;
&lt;br /&gt;
=== H ===&lt;br /&gt;
* [[Handball Club]]&lt;br /&gt;
* [[Heart Club]]&lt;br /&gt;
* [[History Graduate Students&#039; Association]]&lt;br /&gt;
* [[Human Resources Management Club]]&lt;br /&gt;
* [[Humanitarian Society]]&lt;br /&gt;
&lt;br /&gt;
=== I ===&lt;br /&gt;
* [[Improv Theatre Society]]&lt;br /&gt;
* [[Integrated Engineering Club]]&lt;br /&gt;
* [[Integrated Sciences Student Association]]&lt;br /&gt;
* [[Inter Varsity Christian Fellowship]]&lt;br /&gt;
* [[International Business Club]]&lt;br /&gt;
* [[International Relations Students&#039; Association]]&lt;br /&gt;
* [[International Students&#039; Association]]&lt;br /&gt;
* [[Ismaili Students&#039; Association]]&lt;br /&gt;
* [[Israel Awareness Club]]&lt;br /&gt;
* [[Italian Club]]&lt;br /&gt;
&lt;br /&gt;
=== J ===&lt;br /&gt;
* [[Japan Association]]&lt;br /&gt;
* [[Jewish Students&#039; Association (Hillel)]]&lt;br /&gt;
* [[Journalists for Human Rights]]&lt;br /&gt;
&lt;br /&gt;
=== K ===&lt;br /&gt;
* [[Kababayan: Filipino Students&#039; Association]]&lt;br /&gt;
* [[Kendo Club]]&lt;br /&gt;
* [[Kids Help Phone UBC]]&lt;br /&gt;
* [[Korean Campus Mission]]&lt;br /&gt;
* [[Korean Intercollegiate Student Society of UBC]]&lt;br /&gt;
* [[Kung Fu Association]]&lt;br /&gt;
&lt;br /&gt;
=== L ===&lt;br /&gt;
* [[Language Enthusiasm]]&lt;br /&gt;
* [[Latin Dance Passion]]&lt;br /&gt;
* [[Le Club Français (UBC French Club)]]&lt;br /&gt;
* [[Librarians Without Borders]]&lt;br /&gt;
* [[Lifeline]]&lt;br /&gt;
* [[Literature Etc.]]&lt;br /&gt;
* [[Love Your Neighbour of UBC]]&lt;br /&gt;
&lt;br /&gt;
=== M ===&lt;br /&gt;
* [[Mahjong Club]]&lt;br /&gt;
* [[Marketing Club]]&lt;br /&gt;
* [[Mechanical Engineering Club]]&lt;br /&gt;
* [[Médecins Sans Frontières UBC]]&lt;br /&gt;
* [[Medical Informatics Association]]&lt;br /&gt;
* [[Meditation Community]]&lt;br /&gt;
* [[Men&#039;s &amp;amp; Women&#039;s Wrestling Club]]&lt;br /&gt;
* [[Mental Health Awareness Club]]&lt;br /&gt;
* [[Microbiology &amp;amp; Immunology Students&#039; Association]]&lt;br /&gt;
* [[Motorcycling Club]]&lt;br /&gt;
* [[Muslim Students&#039; Association]]&lt;br /&gt;
&lt;br /&gt;
=== N ===&lt;br /&gt;
* [[Navigators]]&lt;br /&gt;
* [[New Democratic Party Club of UBC]]&lt;br /&gt;
* [[New Taiwanese Generation]]&lt;br /&gt;
* [[Newman Catholic Church]]&lt;br /&gt;
&lt;br /&gt;
=== O ===&lt;br /&gt;
* [[Organ Donation Club]]&lt;br /&gt;
* [[Origami/Paper Folding Club]]&lt;br /&gt;
* [[Oxfam UBC]]&lt;br /&gt;
&lt;br /&gt;
=== P ===&lt;br /&gt;
* [[Pakistan Students&#039; Association]]&lt;br /&gt;
* [[Peace &amp;amp; Love UBC]]&lt;br /&gt;
* [[Persian Club]]&lt;br /&gt;
* [[Perspectives]]&lt;br /&gt;
* [[Philosophy Sstudents&#039; Association]]&lt;br /&gt;
* [[Photographic Society]]&lt;br /&gt;
* [[Phrateres UBC, Theta Chapter]]&lt;br /&gt;
* [[Players&#039; Club]]&lt;br /&gt;
* [[Polish Students&#039; Society]]&lt;br /&gt;
* [[Political Science Students&#039; Association]]&lt;br /&gt;
* [[PostSecret Club]]&lt;br /&gt;
* [[Pottery Club]]&lt;br /&gt;
* [[Pre-Dental Society]]&lt;br /&gt;
* [[Pre-Law Society]]&lt;br /&gt;
* [[Pre-Medical Society]]&lt;br /&gt;
* [[Pre-Optometry Club]]&lt;br /&gt;
* [[Pre-Pharmacy Club]]&lt;br /&gt;
* [[Psychology Students&#039; Association]]&lt;br /&gt;
&lt;br /&gt;
=== R ===&lt;br /&gt;
* [[Radical Beer Faction]]&lt;br /&gt;
* [[Real Estate Club]]&lt;br /&gt;
* [[Reality Club]]&lt;br /&gt;
* [[Red Cross International]]&lt;br /&gt;
* [[Right to Play at UBC]]&lt;br /&gt;
* [[Ringette Team]]&lt;br /&gt;
* [[Rotaract Club]]&lt;br /&gt;
* [[Rovers of UBC - Scouts Canada]]&lt;br /&gt;
&lt;br /&gt;
=== S ===&lt;br /&gt;
* [[Safety and First Aid Club]]&lt;br /&gt;
* [[Sailing Club]]&lt;br /&gt;
* [[Salsa Club]]&lt;br /&gt;
* [[Scandinavian &amp;amp; Nordic Cultural Association]]&lt;br /&gt;
* [[Science Fiction Society]]&lt;br /&gt;
* [[Science One Survivors]]&lt;br /&gt;
* [[Serbian Student Association]]&lt;br /&gt;
* [[Seri Malaysia Club]]&lt;br /&gt;
* [[Shito-Ryu Seiko-Kai Karate]]&lt;br /&gt;
* [[Shotokan Karate]]&lt;br /&gt;
* [[Sikh Students&#039; Association]]&lt;br /&gt;
* [[Singapore Raffles Club]]&lt;br /&gt;
* [[Single Parents on Campus]]&lt;br /&gt;
* [[Skateboard Club]]&lt;br /&gt;
* [[Ski &amp;amp; Board Club]]&lt;br /&gt;
* [[Smiling Over Sickness]]&lt;br /&gt;
* [[Sociology Student Association]]&lt;br /&gt;
* [[Speech &amp;amp; Linguistics Student Association]]&lt;br /&gt;
* [[Sports Car Club]]&lt;br /&gt;
* [[Sprouts]]&lt;br /&gt;
* [[Sri Lanka Society]]&lt;br /&gt;
* [[STAND UBC]]&lt;br /&gt;
* [[Storm Club]]&lt;br /&gt;
* [[Students for a Democratic Society]]&lt;br /&gt;
* [[Students for a Free Tibet]]&lt;br /&gt;
* [[Students for Equality &amp;amp; Freedom in the Middle East]]&lt;br /&gt;
* [[Students in Free Enterprise]]&lt;br /&gt;
* [[Students Taking Initiatives]]&lt;br /&gt;
* [[Surf Club]]&lt;br /&gt;
* [[Swing Kids]]&lt;br /&gt;
* [[Synchronized Swim Club]]&lt;br /&gt;
&lt;br /&gt;
=== T ===&lt;br /&gt;
* [[Table Tennis Club]]&lt;br /&gt;
* [[Taekwondo Club]]&lt;br /&gt;
* [[Taiwan Association]]&lt;br /&gt;
* [[Tax Assistance Clinic]]&lt;br /&gt;
* [[Tennis Club of UBC]]&lt;br /&gt;
* [[Thai Aiyara Club]]&lt;br /&gt;
* [[Transportation &amp;amp; Logistics Club]]&lt;br /&gt;
* [[Triathlon Club]]&lt;br /&gt;
* [[Trivia Club]]&lt;br /&gt;
* [[Turkish Students&#039; Society]]&lt;br /&gt;
&lt;br /&gt;
=== U ===&lt;br /&gt;
* [http://ubcimprov.com/ UBCimprov]&lt;br /&gt;
* [[United Nations Children Emergency Fund (UNICEF UBC)]]&lt;br /&gt;
* [[Universities Allied for Essential Medicines]]&lt;br /&gt;
* [[University Christian Ministry]]&lt;br /&gt;
* [[UTSAV: The Indian Students&#039; Association]]&lt;br /&gt;
* [[uVote: Political Awareness Club]]&lt;br /&gt;
&lt;br /&gt;
=== V ===&lt;br /&gt;
* [[Vancouver Students Entrepreneurship Association]]&lt;br /&gt;
* [[Varsity Outdoor Club]]&lt;br /&gt;
* [[Venture Capital &amp;amp; Private Equity Club]]&lt;br /&gt;
* [[Vietnamese Students&#039; Society]]&lt;br /&gt;
* [[Visual Arts Students&#039; Association]]&lt;br /&gt;
&lt;br /&gt;
=== W ===&lt;br /&gt;
* [[Walter Gage Toastmasters]]&lt;br /&gt;
* [[Wargamers&#039; Society]]&lt;br /&gt;
* [[Water Polo Club]]&lt;br /&gt;
* [[Willing Hearts International Association of UBC]]&lt;br /&gt;
* [[Wine Tasting Club]]&lt;br /&gt;
* [[Wing Chun Internal Kung Fu Club]]&lt;br /&gt;
* [[Women Involved in Legislative Leadership Association]]&lt;br /&gt;
* [[World University Service of Canada UBC]]&lt;br /&gt;
* [[World Vision Club]]&lt;br /&gt;
&lt;br /&gt;
=== Y ===&lt;br /&gt;
* [[Yoga Club]]&lt;br /&gt;
* [[Young Women in Business]]&lt;br /&gt;
* [[YOURS Student Association]]&lt;br /&gt;
* [[Youth Outreach Club]]&lt;br /&gt;
&lt;br /&gt;
==Non-AMS Affiliated Clubs==&lt;br /&gt;
* [[Math Club]]&lt;br /&gt;
&lt;br /&gt;
==See also==&lt;br /&gt;
* [[List of Okanagan Clubs]]&lt;br /&gt;
* [[List of Constituencies]]&lt;br /&gt;
[[Category:Student Groups]]&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=User:Ezhao&amp;diff=94948</id>
		<title>User:Ezhao</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=User:Ezhao&amp;diff=94948"/>
		<updated>2011-05-18T00:19:25Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
==Welcome to my page==&lt;br /&gt;
&lt;br /&gt;
My name is Eric Zhao, and I am a third year student in Physiology with a proposed minor in Physics.&lt;br /&gt;
&lt;br /&gt;
I am excited to be a part of UBC wiki and hope that it will be a wonderful shared resource for all.  I hope that teachers get actively involved, as I see this as an opportunity to write out unhelpful and expensive textbooks in favour of streamlined, relevant, and open-source learning resources.&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=User:Ezhao&amp;diff=94947</id>
		<title>User:Ezhao</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=User:Ezhao&amp;diff=94947"/>
		<updated>2011-05-18T00:17:37Z</updated>

		<summary type="html">&lt;p&gt;Ezhao: Initiated my user page&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;My name is Eric Zhao, and I am a third year student in Physiology with a proposed minor in Physics.&lt;br /&gt;
&lt;br /&gt;
I am excited to be a part of UBC wiki and hope that it will be a wonderful shared resource for all.  I hope that teachers get actively involved, as I see this as an opportunity to write out unhelpful and expensive textbooks in favour of streamlined, relevant, and open-source learning resources.&lt;/div&gt;</summary>
		<author><name>Ezhao</name></author>
	</entry>
</feed>