Experiment / E0O98VC0EWhole-Genome STARR-seq (WHG-STARR-seq)

Whole-genome STARR-seq enhancer activity in CHOK1SV GS-KO cells

Comprehensive Mapping of Functional Enhancers in Chinese Hamster Ovary Cells

A genome-wide episomal STARR-seq plasmid library of randomly sheared CHO genomic DNA fragments (approximately 800 bp) was tested in CHOK1SV GS-KO cells with a core murine CMV promoter. Two biological replicates were harvested 6 h after transfection, and the table contains the authors' final 13,395 RNA-over-input enhancer peak calls with genomic sequences and annotations.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated (6 h post-transfection)

Episomal plasmid STARR-seq used a synthesized screening vector with a core mCMV promoter upstream of an intron, truncated maxGFP, and the genomic-fragment cloning site. CHO genomic DNA was sonicated and size-selected at approximately 800 bp (700-900 bp); 200 million cells were transfected per screen with 2 micrograms of library and PEI, and RNA was harvested 6 h later. Reporter RNA and plasmid-input libraries were sequenced on an Illumina NextSeq with paired-end 75-cycle reads. Activity is the source STARR-seq RNA/input enrichment and its log2 transform.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (16 of 16)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 16 definitions
element_id
Unique STARR-seq peak identifier from the authors' Table S3.
source_rank
Source rank column N in the authors' activity-ranked peak table.
chromosome
Chromosome name reported for the CHO assembly sequence.
chromosome_accession
NCBI RefSeq chromosome or contig accession used for the genomic coordinates.
start_position
Source start coordinate for the STARR-seq peak.
end_position
Source end coordinate for the STARR-seq peak.
sequence
Genomic insert sequence (FASTA) tested in the whole-genome STARR-seq library.
sequence_length
Length of the supplied genomic insert sequence in base pairs.
starr_fc
Source STARR-seq fold change, representing RNA-over-input enrichment.
log2_activity
Source log2-transformed STARR-seq RNA-over-input enrichment.
p_value
Source STARRPeaker p-value for the RNA-versus-input peak call.
nearest_promoter_id
Nearest promoter transcript or feature identifier assigned by HOMER.
nearest_gene
Nearest gene name or gene feature assigned by HOMER.
genomic_annotation
HOMER genomic feature annotation for the peak.
distance_to_tss
Source distance from the peak to the nearest transcription start site in base pairs.
strength_class
Derived source-style strength class: strong for log2 activity >= 4.0, weak for <= 3.1, and intermediate between those thresholds.

Quality control

The authors mapped paired-end input and RNA reads with Bowtie2 v2.5.3, retained properly paired uniquely mapped reads with SAMtools v1.19.1 (view -F 3852 -f 2 -q 40), and normalized genome coverage to RPKM. The two biological replicates correlated at R^2 = 0.94 and were combined. STARRPeaker was run against input; the authors filtered 71,493 candidate calls for p < 1e-5, enrichment over input > 2.5, and regular chromosomes 1-10 and X, yielding the 13,395-row final supplementary Table S3. The packaged table retains every row in that author-supplied final table after validating finite numeric activity fields, valid DNA sequence characters, and unique element IDs.

Curation notes

The source table is the authors' post-filtered final whole-genome STARR-seq enhancer set, not a raw read-level table. Coordinates and chromosome accessions are preserved exactly as supplied; sequence length equals end_position - start_position for these source coordinates. The paper also reports a 6-versus-24 h harvest-time pilot and YY1 motif-mutagenesis STARR-seq, but the supplied supporting files provide those results graphically rather than as machine-readable numeric rows, so they are not represented as additional experiments. The individual GFP/SEAP validation table is retained in raw_data but is not labeled as an MPRA experiment.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.