Experiment / E1CC8N86WStandard STARR-seq

FACS-assisted whole-genome STARR-seq in Nipponbare rice protoplasts

Genome‐wide prediction of activating regulatory elements in rice by combining STARR‐seq with FACS

A whole-genome sheared Nipponbare genomic library was cloned into an episomal self-transcribing reporter and transfected into rice stem and sheath protoplasts. GFP-positive cells were isolated by FACS before ARE RNA/cDNA sequencing, with matched plasmid-input libraries in two independent replicates; the packaged table contains the authors' 1,818 final high-confidence genomic regulatory elements.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

FACS-assisted conventional STARR-seq using an episomal vector containing a minimal cauliflower mosaic virus 35S promoter, the first intron of castor bean catalase (Cat1), EGFP, a multiple cloning site, and a NOS terminator. Nipponbare genomic DNA was sheared mainly to 700–1500 bp, cloned into the reporter 3′ UTR, and delivered to stem/sheath protoplasts by PEG transfection. GFP-positive cells were sorted with a 488-nm FACS Aria III workflow; ARE RNA/cDNA and matched plasmid-input libraries were sequenced as paired-end 150-bp reads on Illumina HiSeq X Ten for R1 and R2. The authors compared ARE treatment with input using Bowtie2 mapping to IRGSP1.0 and MACS2 peak calling.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (41 of 41)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 41 definitions
element_id
Stable genomic element identifier formed from the source chromosome and coordinate interval.
source_fasta_header
Original G1/G2 FASTA header from Data S2, including the authors' GC-content group label.
chromosome
Rice chromosome reported by the source workbook.
start
Start coordinate of the supplied element interval, preserved from Data S2.
end
End coordinate of the supplied element interval, preserved from Data S2.
element_length_bp
Length of the supplied element sequence in base pairs; equals end - start + 1 for these records.
sequence
Nucleotide sequence of the tested genomic insert from the G1 or G2 sheet.
gc_content_percent
GC percentage calculated from the supplied element sequence.
gc_group
Authors' sequence cluster: G1 or G2, based on GC-content pattern.
gc_subgroup
Authors' G1 positional subgroup, G1a or G1b; blank for G2.
author_retained_high_confidence
TRUE for records included in the authors' final high-confidence G1/G2 element set after FE > 20 screening.
replicate_support_count
Number of independent source peak sheets (R1 and R2) with an overlapping call for the element.
replicate_1_peak_present
TRUE when the element interval overlaps a Replicate 1 source peak call; otherwise FALSE.
replicate_1_peak_start
Start coordinate of the selected overlapping Replicate 1 peak call; blank when not detected.
replicate_1_peak_end
End coordinate of the selected overlapping Replicate 1 peak call; blank when not detected.
replicate_1_peak_length_bp
Length of the selected Replicate 1 source peak call.
replicate_1_abs_summit
Absolute summit coordinate of the selected Replicate 1 MACS2 peak.
replicate_1_pileup
MACS2 pileup value for the selected Replicate 1 peak.
replicate_1_log10_pvalue
MACS2 -log10(p-value) for the selected Replicate 1 peak.
replicate_1_fold_enrichment
MACS2 fold enrichment over input for the selected Replicate 1 peak.
replicate_1_log2_fold_enrichment
Derived log2 of Replicate 1 fold enrichment over input.
replicate_1_log10_qvalue
MACS2 -log10(q-value) for the selected Replicate 1 peak.
replicate_1_source_cutoff_pass
TRUE when the selected Replicate 1 source peak has fold enrichment > 20, the paper's final screening cutoff.
replicate_1_match_overlap_bp
Number of base pairs overlapping between the element interval and the selected Replicate 1 source peak.
replicate_1_match_overlap_fraction
Overlap divided by the shorter of the element and selected Replicate 1 peak intervals.
replicate_1_match_count
Number of Replicate 1 source peak calls overlapping the element interval before selecting the best overlap.
replicate_2_peak_present
TRUE when the element interval overlaps a Replicate 2 source peak call; otherwise FALSE.
replicate_2_peak_start
Start coordinate of the selected overlapping Replicate 2 peak call; blank when not detected.
replicate_2_peak_end
End coordinate of the selected overlapping Replicate 2 peak call; blank when not detected.
replicate_2_peak_length_bp
Length of the selected Replicate 2 source peak call.
replicate_2_abs_summit
Absolute summit coordinate of the selected Replicate 2 MACS2 peak.
replicate_2_pileup
MACS2 pileup value for the selected Replicate 2 peak.
replicate_2_log10_pvalue
MACS2 -log10(p-value) for the selected Replicate 2 peak.
replicate_2_fold_enrichment
MACS2 fold enrichment over input for the selected Replicate 2 peak.
replicate_2_log2_fold_enrichment
Derived log2 of Replicate 2 fold enrichment over input.
replicate_2_log10_qvalue
MACS2 -log10(q-value) for the selected Replicate 2 peak.
replicate_2_source_cutoff_pass
TRUE when the selected Replicate 2 source peak has fold enrichment > 20, the paper's final screening cutoff.
replicate_2_match_overlap_bp
Number of base pairs overlapping between the element interval and the selected Replicate 2 source peak.
replicate_2_match_overlap_fraction
Overlap divided by the shorter of the element and selected Replicate 2 peak intervals.
replicate_2_match_count
Number of Replicate 2 source peak calls overlapping the element interval before selecting the best overlap.
qc_status
Packaging QC result; PASS indicates a valid nucleotide sequence and a coordinate-length consistency check.

Quality control

The authors aligned reads to the Nipponbare IRGSP1.0 genome with Bowtie2, retained mapped reads passing SAMtools -q 30 -f 2 -F 264, and called ARE-versus-input transcription peaks with MACS2 (-g 4.4e8 -q 0.01 --keep-dup all). They retained duplicates because the library used minimal PCR amplification. For packaging, the authors' final G1/G2 sequence set was used as the element universe; this is the 1,818-element set produced after the reported FE > 20 screening step, rather than all 29,172/26,916 source peak rows. Source calls were joined to this universe by genomic interval overlap, and elements were not discarded when a particular replicate had no overlapping call. Every retained sequence contains only A/C/G/T/N and has a length equal to its supplied inclusive coordinate interval; all 1,818 rows passed package QC. The table reports source peak metrics for each replicate where an overlapping call exists; FASTQ files were not downloaded, so no additional read-level QC was performed.

Curation notes

This is one MPRA experiment with two independent sequencing replicates, not two different biological conditions. The source workbook's Replicate 1 and Replicate 2 sheets contain 29,172 and 26,916 MACS2 calls, respectively; the G1 and G2 sheets contain the authors' 1,818 final sequence records used for downstream analysis. The interval-overlap join reproduces the paper's reported support split of 1,169 elements in both replicates, 372 in R1 only, and 277 in R2 only. The author-retained flag refers to inclusion in the G1/G2 sequence universe; the per-replicate source_cutoff_pass fields instead evaluate the selected joined source peak and can differ when independently represented peak boundaries overlap. G1/G2 and G1a/G1b are sequence-clustering labels, not fluorescence bins or allele labels. The paper's traditional STARR-seq validation constructs and CRISPR disruption assays were not made separate experiments because no comparable genome-wide quantitative MPRA table was supplied for those assays. Rice protoplasts have no resolved Cell Ontology or Cellosaurus identifier in the requested vocabulary, so the biosample is explicitly marked UNMAPPED.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.