Experiment / E4RURO2UQStandard STARR-seq

UMI-STARR-seq validation of 12,000 maize candidate CREs

Precise engineering of gene expression by editing plasticity

A synthetic library of 12,000 candidate 200-bp cis-regulatory elements was cloned into an episomal STARR-seq reporter and transiently transfected into maize leaf protoplasts. Four biological replicates were sequenced as cDNA and plasmid-input libraries; forward and reverse insert orientations were analyzed independently with DESeq2 RNA/input activity scores.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

UMI-STARR-seq in maize protoplasts using an episomal pMD18-T-derived reporter containing a CaMV 35S minimal promoter (-50 to +5 bp), cat-1 intron, GFP coding sequence, and ccdB-containing cloning region. Each 238-mer synthetic oligo contained inverted 19-bp Tn5 mosaic ends and a 200-bp candidate CRE; the reporter was PEG-transfected into protoplasts from second leaves of etiolated seedlings and incubated at 25 °C in the dark for 16 h. Poly(A)+ reporter RNA and re-isolated plasmid DNA were sequenced on an Illumina NovaSeq X Plus with 10-bp UMIs. Forward and reverse insertion orientations were treated as independent measurements.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (16 of 16)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 16 definitions
element_id
Unique processed element identifier combining the source sequence ID and reporter insertion orientation.
sequence_id
Numeric Seq_id from Additional file 7, identifying the source 200-bp candidate CRE.
orientation
Reporter insertion orientation used for the activity measurement: forward or reverse.
chromosome
Chromosome label derived from the source Chr column by adding the chr prefix.
peak_start
Start coordinate of the source candidate CRE interval on the B73 reference genome.
peak_end
End coordinate of the source candidate CRE interval on the B73 reference genome.
gene_id
Associated maize gene identifier reported in the source workbook.
gene_strand
Strand of the associated gene reported in the source workbook; blank for source records without a strand.
contribution_score
Model contribution score reported for the candidate CRE; blank where the source workbook does not report one.
library_sequence
200-bp candidate CRE sequence as reported in the supplementary library table, in its source forward 5-prime-to-3-prime orientation.
reporter_sequence
Sequence in the 5-prime-to-3-prime orientation of the reporter insert; identical to library_sequence for forward rows and reverse-complemented for reverse rows.
group
Candidate-library group from the source workbook: Group 1 through Group 5.
log2_fold_change
Orientation-specific DESeq2 log2 fold-change for cDNA reporter output relative to plasmid DNA input.
adjusted_p_value
Orientation-specific Benjamini-Hochberg adjusted p-value from the DESeq2 RNA/input test.
enhancer_call
TRUE when log2_fold_change > 0 and adjusted_p_value <= 0.05, following the paper's enhancer definition; otherwise FALSE.
qc_pass
TRUE for all rows retained from source records marked Pass quality control = 1.

Quality control

The authors mapped DNA-input and cDNA reads to the 12,000-sequence reference with Bowtie2, retained reads with MAPQ >30, correct length, and no mismatches, collapsed reads by UMI, and retained sequence entries with more than 10 reads in both replicates. The source supplementary table exposes this sequence-level result as Pass quality control. For the processed table, only the 10,680 source sequence IDs marked as QC-passing were retained, and orientation records with missing log2 fold-change or adjusted p-value were excluded. This yields 21,152 orientation-specific records and reproduces the paper's 8,650 enhancer calls using log2 fold-change >0 and adjusted p-value <=0.05.

Curation notes

This package treats the study's UMI-STARR-seq validation as one Standard STARR-seq experiment; the model-training, CRISPR editing, LUC, and tocopherol experiments are not separate MPRA experiments. The paper's narrative describes the library as assessed in duplicate, while the methods and SRA run metadata describe four biological replicates with eight cDNA and eight input sequencing runs (two indexed runs per replicate). Additional file 7 contains 12,000 source sequence rows, of which 1,320 are marked QC=0. The 10,680 retained source rows generate 21,152 complete orientation records after removing 208 missing orientation results, matching the reported analyzed orientation count. The source table has 200-bp sequences for all retained rows; Group 5 control rows may have blank strand and contribution-score fields. Chromosome labels are normalized to chrN, and reverse reporter sequences are derived as reverse complements while preserving the source sequence in library_sequence. No raw FASTQ/SRA read files are included per package scope; PRJNA1211828 run metadata is included in raw_data.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.