Experiment / E5IXX4105Standard STARR-seq

Short pZS*11_4enh fragment library in tobacco leaves

Identification of Plant Enhancers and Their Constituent Elements by STARR-seq in Tobacco Leaves

A Tn5-fragmented pZS*11_4enh library containing four known plant enhancers was inserted upstream of the 35S minimal promoter and assayed in transiently transformed tobacco leaves. This table summarizes the public Fig5D_Rep1 plasmid-input and cDNA barcode counts at unique fragment coordinates after package QC.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated; 2 d after Agrobacterium infiltration under the normal light/dark growth cycle

The pZS*11_4enh plasmid contains the CaMV 35S, pea AB80, wheat Cab-1, and pea rbcS-E9 enhancer regions. Tn5-fragmented plasmid inserts were cloned immediately upstream of the 35S minimal promoter in the STARR-seq reporter, which contained a 15-bp random barcode in the GFP open reading frame. The short-fragment library comprised approximately 5,700 fragments linked to approximately 73,000 barcodes; reporter-mRNA cDNA and plasmid-input barcode frequencies were used to calculate fragment activity. fragment_sequence is reported in the cloned orientation, reverse-complemented for subassembly records marked with strand '-'.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (19 of 19)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 19 definitions
fragment_id
Unique fragment identifier derived from pZS*11_4enh start, stop, and strand coordinates.
plasmid_start_1based
1-based inclusive start coordinate on the supplied pZS*11_4enh plasmid sequence.
plasmid_stop_1based
1-based inclusive stop coordinate on the supplied pZS*11_4enh plasmid sequence.
strand
Orientation of the fragment relative to the supplied plasmid sequence.
fragment_length_bp
Fragment length in base pairs from the subassembly file.
fragment_sequence
Fragment nucleotide sequence in the cloned orientation; reverse-complemented for strand '-'.
library_assembly_count
Median subassembly support count among retained barcodes linked to this fragment.
barcode_count
Number of linked barcodes passing package QC and aggregated for the fragment.
input_count_sum
Sum of plasmid-input barcode counts across retained barcodes.
cdna_count_sum
Sum of recovered reporter-mRNA cDNA barcode counts across retained barcodes.
log2_enrichment_median
Median log2(cDNA barcode frequency / input barcode frequency) across retained barcodes.
log2_enrichment_mean
Mean log2(cDNA barcode frequency / input barcode frequency) across retained barcodes.
log2_enrichment_sd
Sample standard deviation of barcode-level log2 enrichment within the fragment.
source_replicate
Source count-file replicate label.
replicates_available
Number of source biological replicates represented in this packaged table.
replicates_in_paper
Number of biological replicates reported for this experiment in the paper.
source_subassembly
Relative path to the barcode-to-fragment subassembly file used for joining.
source_input_count_file
Relative path to the plasmid-input barcode count file used.
source_cdna_count_file
Relative path to the reporter-mRNA cDNA barcode count file used.

Quality control

The authors filtered barcode counts below 5 and, for the complete study analysis, discarded barcodes present in only one of three biological replicates. Their reported replicate-quality criteria were Spearman correlation of at least 0.6 for individual barcodes and at least 0.7 for aggregated fragments or variants. For this package, the only compact count files released with the analysis repository are Fig5D_Rep1 input and cDNA counts, so the across-replicate presence filter cannot be applied; retained entries instead have a linked subassembly record and both input and cDNA counts of at least 5. Counts were normalized over the joined mapped barcode set, and retained barcode-level enrichments were aggregated by unique start, stop, strand, and length.

Curation notes

The paper performed three biological replicates, but the linked repository exposes compact count outputs for Fig5D Rep1 only; the complete 84-run sequencing release remains available through BioProject PRJNA627258 and is represented by run metadata in raw_data. This package summarizes 75,036 retained mapped barcodes into 5,661 unique fragment coordinate/orientation rows. The supplied GitHub FASTA is 5,282 bp long; fragment sequences and endpoint validation use that file, while the Addgene page describes the plasmid as 5,283 bp. Leaf is a plant anatomical term without a resolved UBERON leaf identifier, so the allowed UNMAPPED biosample fallback is used.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.