Experiment / E2M8HDSESEpisomal Plasmid MPRA

Natural soybean STEs in the EF1a-promoter STEM-seq screen

From Natural Discovery to AI‐Guided Design: A Curated Collection of Compact Enhancers for Crop Engineering

An 80-bp natural short transcriptional enhancer (STE) library was assayed in Glycine max (soybean) using the STEM-seq plasmid reporter with the EF1apro context. The table contains the authors’ two-replicate RNA/DNA-normalized activity measurements and their replicate-combined statistics.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

STEM-seq is a custom plasmid MPRA: each candidate 80-bp sequence was placed upstream of the EF1apro reporter promoter in pSTEM02 and represented by three independent 9-bp barcodes fused to the N-terminus of firefly luciferase. Species-specific introns distinguish reporter cDNA from input plasmid DNA, GFP monitors transformation, and barcode-containing RNA and plasmid amplicons were sequenced by Illumina HiSeq PE150. Activity is the RNA/plasmid fold change normalized against the low-TPM control set; two independent biological replicates were performed.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (16 of 16)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 16 definitions
source_record_id
Unique processed row identifier; equals the source element ID except when the source repeats an accession, where a source-record suffix disambiguates the repeated measurements.
element_id
Source STE accession identifier from the STEM-seq library; use source_record_id for a unique processed-row key when the source accession is repeated.
sequence
Tested natural 80-bp sequence in the 5-prime to 3-prime orientation.
sequence_length_bp
Sequence length in base pairs after sequence QC.
origin_gene
Source gene locus associated with the upstream sequence in the supplementary candidate/STE annotation.
candidate_tpm
TPM of the source gene used for candidate selection, from Supplementary Sheet 2.
activity_repeat1
Author-reported RNA/DNA/control-normalized enhancer activity for biological replicate 1 (fold change).
activity_repeat2
Author-reported RNA/DNA/control-normalized enhancer activity for biological replicate 2 (fold change).
activity_mean
Author-reported mean enhancer activity across the two biological replicates (fold change).
zscore_repeat1
Z-score of replicate 1 activity relative to the low-expression/negative-control background.
zscore_repeat2
Z-score of replicate 2 activity relative to the low-expression/negative-control background.
combined_p_value
One-tailed replicate-combined P-value calculated with Fisher’s method.
fdr_bh
Benjamini-Hochberg adjusted P-value reported by the authors.
q_value
Storey q-value reported by the authors.
significant_fdr
Author-reported indicator that the BH-FDR threshold was met (1 = significant, 0 = not significant).
significant_q_value
Author-reported indicator that the q-value threshold was met (1 = significant, 0 = not significant).

Quality control

The authors merged and counted barcode reads, reporting >99% successful read combination, and used two biological replicates. Their statistical QC combined one-tailed replicate P-values with Fisher’s method and applied Benjamini-Hochberg FDR and Storey q-value thresholds of <0.05; the source flags are retained in the table (76 BH-significant and 76 q-value-significant rows). For package-level QC, rows were retained only when they had a non-empty element ID, a canonical 80-bp A/C/G/T sequence, and finite values for both replicate activities, the mean, both Z-scores, combined P-value, FDR, q-value, and significance flags. The downloadable source had 2180 data records; 1975 passed these checks and 205 sequence-less or otherwise incomplete records were omitted. Non-significant measured elements are retained as biologically informative low-activity outcomes rather than being treated as failed QC.

Curation notes

The detailed source is raw_data/Sup_Dataset 3.xlsx, sheet soybean_EF1a, joined to raw_data/ADVS-13-e16600-s007.xlsx for source-gene/TPM annotations. The article narrative reports 6904 functional natural STEs with species counts of 2581 maize, 452 wheat, 1896 tomato, and 1975 soybean, whereas the downloadable detailed sheets provide 1975 valid sequence-bearing rows for this condition. This package follows the detailed supplementary measurements and preserves their activity/significance flags. 0 retained rows had no matching source-gene annotation. The table contains screen-level STEM-seq measurements, not the separate dual-luciferase validation assay.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.