MotifView
To guide the user through the different steps of using ChiPSummitDB, the POU2F2 motif is used as an example because it has a relatively small set of overlapping factors.
Click on the “Search views” menu and then click the “MotifView” button.
In the “Set a motif” dropdown box, select POU2F2.
Click on the “Go to MotifView” button to update the page.
After updating the page, a scatterplot for the POU2F2 motif can be seen. Filter out the experiments, which overlap with POU2F2 less than 500 times. To do this, change the default value (100) to 500 in the “Minimum overlap number between motifs and peaks of experiment” box. Click on the “Refresh page” button to update the scatterplot. Only 28 scatter should now be visible. This means that the identified POU2F2 binding sites overlap with peaks from 28 ChIP-seq experiments. However, on the right-hand side of the page, only 20 TFs are displayed and labeled with different colors. This indicates that the 28 scatters on the plot are from 28 ChIP-seq experiments (but some experiments have the same protein target > 20 different protein). Thus, certain TFs are presented in multiple experiments as indicated by the number in the front of the TF’s name (for example “5 P300”).
For simplicity, investigate the TFs individually. To hide all scatters click on the “Hide all scatters” button, which results in an empty graph. To undo click on the “Show all scatters” button. The user can also display individual data points by selecting the TF’s name.
The legend in the right-hand side window shows that data from two POU2F2 ChIP-seq experiments are available. After hiding all scatters, click on the purple square next to POU2F2 and two scatters will appear in the graph area.
By moving the cursor over any of the two scatters, detailed information about the experiment will be displayed at the tool tip. The description indicates, that the data is from ChIP-seq experiments in which the lymphoblastoid cell lines, GM10847 and GM12878, were used. The number of elements is 4179 for cell line GM12878, indicating 4179 POU2F2-binding motifs in the genome and peaks of the sequenced reads (tags) of TF POU2F2 overlap. The tooltip also shows the average distance between the peak summits, i.e. the maxima position of protein covered regions, and the motif centers, i.e. the central base pair of the protein binding DNA motifs, which are fixed reference points in the genome. The standard deviation value for the average distance is shown also. By comparing the average distances from tooltips for different scatters, the position preferences of the protein can be elucidated. Regarding the position values of 1.07 and 1.47 (for the two POU2F2 experiments) in the example, we can conclude that POU2F2 covers the DNA downstream of the motif to which it binds, and also binds to its own regulatory region. The summit positions have relatively low standard deviations (< 20). We postulate that direct protein-DNA binding is more stable and less mobile than indirect contact, and this leads to the relatively low standard deviation values for these binding sites. Here are two more examples for the above postulate. When the two POU2F2 scatters are displayed, click on the colored square next to P300. Five triangles will appear on the plot. It shows that P300 binding preferences are located on either side of the zero reference point and their SD is higher than POU2F2’s. This suggests that P300 binds indirectly to the DNA via POU2F2. Next click on the colored square next to EB1, which will bring up two more scatters. Although the binding position of EB1 overlaps with that of POU2F2, EB1 has a different binding motif. Thus, we hypothesize that both P300 and EB1 bind POU2F2 and the complex binds to the binding site of POU2F2.
Interoperability between the different views
The MotifView provides a complex picture of the transcription factor binding sites and the overlapping ChIP-seq read peaks in an interactive graph. Scatter on the diagram, represents one particular ChIP-seq experiment and can be selected by clicking on it. A maximum of three scatters can be selected.
After selecting scatters, certain information (cell line, antibody, and element number) about the corresponding experiment will appear in the three boxes under each other below the graph. To view detailed information there are the following options:
Clicking on the “to paired shifts” button, will open a new window with the distance distribution chart of the summit positions of the selected experiments, compared to that motif, which was selected in the previous step.
To browse the genomic locations of the peak-motif co-occurrences in the genome browser, click on the “to the GenomeView” button.
To see the number of common and specific peak-motif pairs between the selected experiments, click on the “to venn diagram” button.
Furthermore, you can download the intersection of the corresponding motif and the last selected experiment by clicking on the “download last selected button”. The selected experiment list can be modified by using the “Clear all selected” button. This will remove all of the selections.
Changing the display of data
The middle and the right-hand side panel of options stand for the modification of the data display.
In the middle, the first two buttons allows the user to the Y value, which will toggle between two states: standard deviation and element number. For further information please check the Help menu. The default setting is the standard deviation.
Click on the “element number” to display the motif- experiment peak overlap number as a Y value.
The page will update and reorder the scatters of experiments depending on the overlap number between ChIP-seq peaks and POU2F2 motifs. Two POU2F2 ChIP-seq data have the highest overlap number of ~4000 POU2F2 binding sites. To find detailed information hover over one of the scatter, which will show the accurate element number and other information. The POU2F2 ChIP-seq derived from GM10847 cell line has 3851 peaks, which co-located with POU2F2 motifs.
The unification step averages the average summit positions, the average standard deviations, and the element numbers obtained from different experiments. To display the unification calculations, the user can choose between “average standard deviation vs. average of average positions” or “average element numbers vs. average of average positions” buttons.
Click on the “average standard deviation vs average of average positions” button.
This step will combine the values of ChIP-seq experiments (with similar antibodies) and create a new plot, where the y value is the average standard deviations.
Only one dot corresponds to each antibody, which shows the average values of its corresponding experiments.
The buttons on the right-hand side modify the order and the content of labels.
The first two buttons sort the antibodies according to their names (alphabetically) or the number of experiments.
Choosing the last two buttons will modify the labels. The antibody names will be replaced with cell type names. Following the logic of the previous buttons, the experiment can be sorted by name of the cell line or the number of their occurrences.
PairShiftView
To introduce this tool, the relation between the CTCF motif and its corresponding summits from two experiments from the MCF7 cell line will be shown as an example.
Move the cursor over the “Search views” menu and click on the “PairShiftView “button.
Select the CTCF motif under “Set a motif”.
Click on the first upper left (red) dropdown box and select the “MCF7” cell type from the list. Click on the second box from the left in the row, and select CTCF from the list. In the third red box, select the name of the following experiment: "hs_BreastAdenocarcinoma_MCF7_cancer_CTCF_SRX1091824”.
In the next (blue) row of boxes, repeat the previous steps, but select RAD21 as the antibody and “hs_BreastAdenocarcinoma_MCF7_cancer_RAD21_ERX004452” as the experiment name.
After setting the parameters, click on the “Refresh Page” button.
The page will refresh in a new window and a diagram will be displayed. The color of the curves corresponds to the colors of the boxes used for settings. If the curves do not fit inside the diagram, then adjust the minimum and maximum values of the axes below the graph.
The page will be updated and the modified diagram will be displayed.
On the plot, the red curve (CTCF) is shifted towards the left-hand side of the center of the CTCF motif, with a peak around position -5. In contrast, the RAD21 peak is shifted to the 3’ direction of the motif and has a local maxima around position 15. Based on these observations, it can be assumed that the fine positional shifts that may exist between the contact points of cohesin proteins (CTCF, RAD21, SMC1/3, and STAG1/2) might reflect the 3D position of the components within the complex.
Explanation: Since CTCF is the only known specific DNA binder among the components of the CTCF/cohesin complex, we expected that the corresponding ChIP-seq peaks will point to the same position with respect to CTCF motif. In contrast, the fact that we can observe a positional shift suggests that RAD21 proteins occupy conserved – relatively fixed – positions that are close enough to the DNA so as to allow DNA-protein crosslinks to form during the ChIP-seq procedure.
The JARID1B is known as a histone demethylase enzyme. A high fraction of JARID1B peaks overlap with CTCF binding sites in basal breast cancer cells (Yamamoto et al., 2014). The relationship between these two factors has been investigated, as well as their relative effects on each other. According to the knock-down experiments, they found clear evidence for CTCF-JARID1B interactions, which suggests that the two proteins are present in the same complex.
When JARID1B and “hs_BreastAdenocarcinoma_MCF7_cancer_JARID1B_SRX265412” are selected in the third row of green boxes, a flat and broad green curve for the third experiment appears in “PairShiftView”, indicating frequent CTCF-JARID1B co-appearance. Consequently, the standard deviation of the positions for the JARID1B curve is higher (20.55), than that of the other two curves. This indicates that the interaction between the CTCF‑binding motif and the JARID1B protein is very likely not direct, the protein occupies the binding sites via several other proteins.
Use the second panel, if you want to start the experiment selection with the name of the antibody. This panel is similar to the previous with a slight difference, the “cell type “ and “antibody” columns are switched.
dbSNPView
Users can find motifs which overlap with known nucleotide variations. The database contains dbSNP entries. This tutorial will demonstrate how to perform this action.
Click on the “Search views” menu and then click the “dbSNPView” button.
There are two ways to find overlapping SNPs. The first one is to specify the dbSNP ID itself. Write “rs1193185173” in the dbSNP search box, then click “Send”.
Now a chromosomic view is displayed with all the SNPs and motifs. You can click on the motif logo and the browser immediately goes to the MotifView. You can also click on the SNPs to see details of the selected variation in a new window. You can also switch off SNPs not overlapping with any motifs.
This view is limited to 100 bp only. The user may want to see a larger genomic landscape. To perform this action, set the chromosome to 10, Start position to 711630, and End position to 711888. Remove the text from the dbSNP box. If dbSNP is set, the website will not set the other fields. Remark: the genomic region cannot be larger than 1000 bp. If every setting is correct, click “Send”.
The user can now see all the motifs and variations in the region. Clicking on the motifs will switch to the MotifView. The overlapping SNPs are red. Clicking on them has a similar effect as if the user specified the SNP id in the website. The non-overlapping SNPs are blue and clicking on them will display details about the SNP.
ExperimentView
The processed ChIP-seq experiment’s attributes can be browsed in this view.
Move the cursor over the “Search views” menu and click on the “ExperimentView“ button.
You can select a given experiment from the list using the dropdown boxes, which are similar to the PairShiftView panels. Click on the left dropdown box and select the “MCF7” cell type from the list. Click on the second box from the left in the row, and select CTCF from the list. In the third box click on the “hs_BreastAdenocarcinoma_MCF7_cancer_CTCF_SRX1091824” experiment.
After the page refreshes, you can read the attributes of each experiment. The following attributes can be found:
Experiment name: Name of ChIP-seq data according to our nomenclature: “Organism (hs- Homo sapiens) _ “Tissue name” _ “Cell type” _ “Cell line type” _ “Antibody name” _ “SRA ID”
Antibody: name of ChIP-seq target protein.
Cell line: Name or code of cell type
Sra ftp link: Link to directly download raw SRA file in .sra format.
Homer de novo motif: The htp report of the homer de novo motif scan under corresponding peak regions.
GenomeView: link to browse the peak region in Jbrowse genome viewer.
Number of peaks: Number of peaks from the HOMER peak calling analysis.
Link to motif view if antibody and consensus motif is the same: This link can be used for transcription factors, which have identified JASPAR CORE consensus motif set. This link navigates to the corresponding motif’s MotifView.
Number of reads: Raw sequence data reads
SRX search: link to the SRA report about the experiment.
Overlapping motifs: List of motifs which are occupied by peaks for a given ChIP-seq. The second column represents the number of co-occurrences. The motif names are hyperlinks which direct the user to the MotifView.
There are two ways to use the GenomeView. After finding some interesting experiment in MotifView you can export them into GenomeView to visualize your finds or you can open it directly and use an “à la carte” system to select what you would like to see. If you chose to get to GenomeView from MotifView you can get more specific information and also use the “à la carte” mode, however if you do not need the experiment:consensus motif specific peak data you can head straight to GenomeView from the Search views menu.
For this tutorial we are going to start with, how to use GenomeView via Motif View first. If you would like you can skip to the description of the “à la carte” system, as that is sufficient if you want to use GenomeView directly. To do so please go to step II/1.
I. Importing tracks from MotifView
To guide through the different steps of using ChIPSummitDB the POU2F2 motif was used as example because it has a relatively small set of overlapping factors. We will start just like in the MotifView tutorial.
I. / 1. Click on the “Search views” menu and then click the “MotifView” button.
I. / 2. In the “Set a motif” dropdown box select POU2F2 and sat “Minimum overlap number between motifs and peaks of experiment” to 500. Than click on “Go to MotifView”
I. / 3. After updating the page, a scatterplot for the POU2F2 motif can be seen. For further description on what the scatterplot shows please refer the MotifView tutorial. Please select the same experiments as seen at the “Interoperability between the different views” and export them in to the genomeView. This is the point where we diverge from the MotifView tutorial.
For now, please click on the 3 peak-sets shown with red arrows (I.). You can notice that after you have clicked on a peak-set it becomes highlighted and under the plot, some details, from the experiment that you clicked on, will appear. Please click on “to the GenomeView”, shown with black arrow.
I. / 6.
You should see something like the next screenshot. Welcome to GenomeView! This is powered by jbrowse v16.3. If you have any familiarity with it or any other genome visualizer (gbrowse, IGV or IGB, etc..) you will probably navigate it easily. If this is your first time with a genome browser the next paragraph will describe how to navigate.
The genome is represented as the X-axis: 5’ to the left and 3’ to the right (if you zoom in enough the sequences are shown). Data tracks are presented under it, one beneath the other. The name of the tracks are on the top right corner of each track. Tracks are showing information based on genomic position. For a description how to move around and zoom click on “Help” (red arrow) and then on “?General”.
I. / 7.
If you look around you will see the first track under the genome is the gene map from the UCSC (https://genome.ucsc.edu/). Under it is the map of the transcription factor bindig sites, in our case that will be for POU2F2. Under this, you will have 3 track for the 3 selected experiments (the order you see them may vary).For a good example please go to “chr1:157696368..157697135” (You can simply enter it at the top).
Here you can see a transcription factor biding site where all 3 peak-sets are present. The peak is represented as a line, the motif is shown as an orange rectangle, and the orange vertical line shows the summit. The number shows the score of the peak calling. This shows that the summit in the p300 track is a bit to the 5’ compared to the other two.
II. Adding tracks à la carte
II. / 1. Tracks can be added to GenomeView from the “select tracks” menu (red arrow).
II. / 2. After clicking on it, a panel with 2 parts shows up from the left. The main part is the table (red rectangle) which shows all the tracks available in the database. Here you will find motif tracks (consensus motifs of ChIPSummitDB), experiment (all peaks from an experiment) and miscellaneous tracks. Other parts are to filter these. On the top (black arrow) you can enter any text you are looking for in the database. On the left you can use filters to show only the type of tracks you need (Green arrows). As there is an overwhelming amount of experiment tracks you can filter these more precisely. At the yellow rectangle you can filter expreiments based on the antibody and/or cell line that was used.
II. / 3. Examples of all available track types
VennDiagramView
You can display the frequency of common and specific peaks for selected experiments at a consensus motif binding site in the VennDiagramView.
Move the cursor over the “Search views” menu and click on the “VennDiagramView” button.
We will use the CTCF motif as an example. Under “Set a motif”, click on the drop down box and select “CTCF” from the list.
Below the motif selection, click on the first upper left (red) dropdown box and select the “MCF7” cell type from the list. Click on the second box from the left in the row, and select CTCF from the list. In the third red box select the name of the experiment to “hs_BreastAdenocarcinoma_MCF7_cancer_CTCF_SRX1091824”.
In the next (blue) row of boxes, repeat the previous steps, but select RAD21 as the antibody and “hs_BreastAdenocarcinoma_MCF7_cancer_RAD21_ERX004452” as the experiment name.
In the third (green) row, select MC7 as the cell type and JARID1B as the antibody. In the third box, choose “hs_BreastAdenocarcinoma_MCF7_cancer_JARID1B_SRX265412” from the list.
After you set all the parameters, click on the “Refresh Page” button to display the data. The sets represent the overlap between the chosen motif and ChIP-seq experiments. The colors of the sets correspond to the colors of the dropdown boxes. The segments show the number of co-occurences between two or three ChIP-seq experiments at the motif. All segment sizes are revealed on the Venn diagram.
Use the second panel, if you want to start the experiment selection with the name of the antibody. This panel is similar to the previous with a slight difference, the “cell type“ and “antibody” columns are switched.