SIF-seq mouse ES-cell enhancer screen at the Nanog locus
Function-based Identification of Mammalian Enhancers Using Site-Specific IntegrationA randomly sheared mouse BAC spanning Nanog and neighboring genes was tested in mouse ES cells using single-copy Hprt-targeted Venus reporter constructs. Fluorescent sorted and unsorted input populations were sequenced to identify enriched enhancer regions.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Perturbation & assay details
Basal / Untreated
SIF-seq uses approximately 1–1.6 kb genomic fragments cloned next to a minimal hsp68 promoter and Venus YFP reporter, followed by homologous-recombination integration as a single copy at the Hprt locus in Hprt-deficient male E14Tg2a.4 mouse ES cells. YFP-positive FACS material and unsorted input were amplified with universal primers and sequenced on a PacBio RS; activity is represented by normalized fluorescent-versus-input target coverage rather than molecular barcode counts. The library re-identifies three Nanog-region SIF-seq sites described in the paper and includes one additional packaged call that passes the stated region filter.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (27 of 27)
| Row | |||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | |||||||||||||||||||||||||||
| 2 | |||||||||||||||||||||||||||
| 3 | |||||||||||||||||||||||||||
| 4 |
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 27 definitions
- element_id
- Stable identifier for the packaged enriched region.
- library
- SIF-seq library name.
- cell_type
- Reporter-cell condition used for the fluorescent sort.
- reference_genome
- Genome assembly for the reported coordinates.
- chromosome
- Chromosome or contig.
- start
- 1-based inclusive genomic start coordinate.
- end
- 1-based inclusive genomic end coordinate.
- region_length_bp
- Length of the enriched region in base pairs.
- input_qc_pass_alignments
- Total input-library alignments retained by the read-level QC.
- fluorescent_qc_pass_alignments
- Total fluorescent-population alignments retained by the read-level QC.
- input_reads_overlapping_region
- QC-passing input reads whose target alignment overlaps the region.
- fluorescent_reads_overlapping_region
- QC-passing fluorescent-population reads whose target alignment overlaps the region.
- mean_input_coverage
- Mean raw input read-span coverage across the region.
- mean_fluorescent_coverage
- Mean raw fluorescent-population read-span coverage across the region.
- max_input_coverage
- Maximum raw input read-span coverage across the region.
- max_fluorescent_coverage
- Maximum raw fluorescent-population read-span coverage across the region.
- max_log2_activity
- Maximum log2((normalized fluorescent coverage + 5)/(normalized input coverage + 5)) in the region.
- median_log2_activity
- Median per-base normalized fluorescent-versus-input log2 activity score in the region.
- max_log2_position
- 1-based genomic position at the maximum log2 activity score.
- empirical_p
- Empirical p-value from the rounded genome-wide background log2 activity distribution.
- merged_peak_count
- Number of thresholded subpeaks merged into this region within 1,000 bp.
- author_description
- Biological or locus annotation reported by the authors or inferred from the paper figures.
- validation_status
- Individual validation status reported in the paper.
- source_input_run
- SRA run accession for the input library.
- source_fluorescent_run
- SRA run accession for the fluorescent sorted library.
- source_reference
- Packaged reference FASTA used for alignment.
- qc_status
- Whether the row passed the packaged read and enriched-region QC.
Quality control
Primary PacBio alignments were retained with MAPQ >=20, aligned query length >=100 bp, query aligned fraction >=0.50, and NM-derived estimated identity >=0.70. Coverage was normalized as in the authors’ PeakAnalysis_Clean.R script, using a 5-count pseudocount; candidate runs required 20 consecutive bases with log2 activity >=0.585 (1.5-fold), and packaged regions additionally required length >=800 bp and empirical p<0.05 as stated in the paper. Rows failing these filters were excluded.
Curation notes
This experiment is represented by paired public PacBio runs SRR1104060 (input) and SRR1104061 (fluorescent), with no replicate run identified in the public run metadata. The source software exposes a min.region.size parameter of 500 bp but does not enforce it inside the region-calling function; the paper text explicitly describes an 800-bp minimum, which was used for this package. Reference sequence and coordinate assembly are recorded in the table. The library re-identifies three Nanog-region SIF-seq sites described in the paper and includes one additional packaged call that passes the stated region filter.