Experiment / E2HXN5N00Standard STARR-seq

Targeted STARR-seq enhancer length optimization

Comprehensive Mapping of Functional Enhancers in Chinese Hamster Ovary Cells

The top 50 whole-genome STARR-seq peaks were fragmented into overlapping 500-bp synthetic inserts and tested in a targeted STARR-seq library with gene-desert and random-sequence negative controls. The table contains valid targeted fragment sequences and RNA/input activity measurements from the authors' supplementary Table S4.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated (6 h post-transfection)

Equimolar synthetic fragments with sequencing adapters and vector homology arms were cloned into the episomal mCMV-core STARR-seq vector. The library included 500-bp bins centered on or flanking the top 50 whole-genome peaks plus gene-desert and random negative controls; 186 fragments were designed and 159 were successfully synthesized. Targeted screens used 20 million CHOK1SV GS-KO cells per screen, RNA was collected 6 h after transfection, and targeted libraries were sequenced on an Illumina MiSeq with paired-end 75-cycle reads. Bowtie v1.3.1 and GATK/Picard CollectHsMetrics were used to count fragment coverage; RPM-normalized RNA/input fold change is the activity measure.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (13 of 13)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 13 definitions
element_id
Unique targeted fragment identifier from the authors' Table S4.
source_rank
Source row rank column N.
parent_peak_id
Whole-genome STARR-seq peak from which the targeted fragment was designed, or a negative-control identifier.
fragment_coordinates
Source genomic coordinate string for the targeted fragment; blank for random-sequence controls.
sequence
500-bp targeted insert sequence supplied by the authors.
sequence_length
Length of the supplied targeted insert sequence in base pairs.
input_rpm
RPM-normalized plasmid-input library coverage for the fragment.
rna_rpm
RPM-normalized STARR-seq RNA coverage for the fragment.
starr_fc
Source fold change of RNA RPM over input RPM.
log2_activity
Derived log2-transformed RNA/input fold change.
summit_relative_position
Source position of the whole-genome peak summit relative to the targeted fragment; blank for controls.
element_class
peak_fragment for fragments derived from a whole-genome peak, or negative_control for gene-desert/random controls.
activity_call
Derived call: active when starr_fc > 1, inactive_or_not_detected when starr_fc < 1, and neutral when equal to 1.

Quality control

The authors report 159 successfully synthesized fragments from 186 designs and quantified each targeted fragment by custom-reference alignment, RPM normalization, and RNA/input fold change. The supplied Table S4 contains 159 rows, including 149 peak-derived fragments and 10 intentional negative controls. One peak_22_2 row had input RPM = 0 and a published #DIV/0! activity value, so it was excluded from processed_data because its activity is undefined. All other rows had numeric RPM/activity values, valid DNA sequence characters, and a sequence length matching the reported 500 or 501 bp; negative controls were retained because they are valid assay measurements rather than QC failures.

Curation notes

The packaged table has 158 rows after removal of only the undefined-input peak_22_2 measurement. The raw workbook remains unchanged and preserves all 159 source rows. The study's harvesting-time pilot and YY1 mutant STARR-seq assay are not separately packaged because their supporting file contains no machine-readable numeric output table.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.