Experiment / E4X2RSZ1ATargeted Genomic Integration MPRA

Iterative FACS-enriched random intron library RNA MPRA

Sequence-dependent and -independent effects of intron-mediated enhancement learned from thousands of random introns

The integrated random-intron reporter library was sorted into GFP/dTomato-high green and GFP/dTomato-low red populations through three successive FACS stages in two independent trajectories. RNA-seq of the twelve sorted bins provides stage-specific spliced GFP, unspliced GFP, and dTomato counts for the same trusted intron elements.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Iterative FACS selection: top and bottom 10% GFP/dTomato populations over three stages

The same single-copy HILO-RMCE HEK293T A2 dual-reporter library was used. Cells were sorted on a BD FACSAria III using GFP/FITC and dTomato/PE signals; each stage selected approximately the top or bottom 10% along the reporter-ratio diagonal, with green (G1–G3) and red (R1–R3) stages in each of two trajectories. The table contains RNA-seq read counts from sorted bins, rather than per-cell fluorescence values, and carries over author-supplied bulk IME and splicing annotations.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (75 of 75)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 75 definitions
element_id
Stable package identifier for the tested element.
barcode
Trusted reporter barcode associated with the intron sequence.
intron_sequence
Full deposited synthetic intron sequence, including constant splice-site flanks.
random_region_sequence
Internal variable sequence after removing the constant splice-site flanks; uppercase DNA.
intron_length
Length of the full deposited intron sequence in nucleotides.
random_region_length
Length of the variable internal sequence in nucleotides.
polyU3_count
Number of overlapping TTT motifs in the variable region; DNA T corresponds to RNA U.
polyU4_count
Number of overlapping TTTT motifs in the variable region; DNA T corresponds to RNA U.
polyU5_count
Number of overlapping TTTTT motifs in the variable region; DNA T corresponds to RNA U.
random_region_gc_fraction
Fraction of variable-region bases that are G or C.
ime_score_log2fc
Authors' DESeq2/Ashr intron-mediated-enhancement log2 fold-change estimate normalized to intronless controls.
deseq_pval
Author-supplied DESeq significance value from supplementary Table S4; blank where the source is blank.
facs_set
Author-supplied iterative-FACS set label: none, green, or red.
median_splicing_efficiency
Author-supplied median fraction of classified GFP reads that were spliced.
gfp_spliced_G1_rep1
Classified spliced GFP reads in G1, trajectory replicate 1.
gfp_unspliced_G1_rep1
Classified unspliced GFP reads in G1, trajectory replicate 1.
dtom_G1_rep1
Classified dTomato reads in G1, trajectory replicate 1.
total_reads_G1_rep1
Sum of dTomato, spliced GFP, and unspliced GFP reads in G1, trajectory replicate 1.
gfp_spliced_G2_rep1
Classified spliced GFP reads in G2, trajectory replicate 1.
gfp_unspliced_G2_rep1
Classified unspliced GFP reads in G2, trajectory replicate 1.
dtom_G2_rep1
Classified dTomato reads in G2, trajectory replicate 1.
total_reads_G2_rep1
Sum of dTomato, spliced GFP, and unspliced GFP reads in G2, trajectory replicate 1.
gfp_spliced_G3_rep1
Classified spliced GFP reads in G3, trajectory replicate 1.
gfp_unspliced_G3_rep1
Classified unspliced GFP reads in G3, trajectory replicate 1.
dtom_G3_rep1
Classified dTomato reads in G3, trajectory replicate 1.
total_reads_G3_rep1
Sum of dTomato, spliced GFP, and unspliced GFP reads in G3, trajectory replicate 1.
gfp_spliced_R1_rep1
Classified spliced GFP reads in R1, trajectory replicate 1.
gfp_unspliced_R1_rep1
Classified unspliced GFP reads in R1, trajectory replicate 1.
dtom_R1_rep1
Classified dTomato reads in R1, trajectory replicate 1.
total_reads_R1_rep1
Sum of dTomato, spliced GFP, and unspliced GFP reads in R1, trajectory replicate 1.
gfp_spliced_R2_rep1
Classified spliced GFP reads in R2, trajectory replicate 1.
gfp_unspliced_R2_rep1
Classified unspliced GFP reads in R2, trajectory replicate 1.
dtom_R2_rep1
Classified dTomato reads in R2, trajectory replicate 1.
total_reads_R2_rep1
Sum of dTomato, spliced GFP, and unspliced GFP reads in R2, trajectory replicate 1.
gfp_spliced_R3_rep1
Classified spliced GFP reads in R3, trajectory replicate 1.
gfp_unspliced_R3_rep1
Classified unspliced GFP reads in R3, trajectory replicate 1.
dtom_R3_rep1
Classified dTomato reads in R3, trajectory replicate 1.
total_reads_R3_rep1
Sum of dTomato, spliced GFP, and unspliced GFP reads in R3, trajectory replicate 1.
green_stage_coverage_rep1
Number of green stages G1-G3 in trajectory 1 with at least 100 total classified reads.
red_stage_coverage_rep1
Number of red stages R1-R3 in trajectory 1 with at least 100 total classified reads.
green_total_reads_rep1
Total classified reads summed over G1-G3 in trajectory 1.
red_total_reads_rep1
Total classified reads summed over R1-R3 in trajectory 1.
trajectory_total_reads_rep1
Total classified reads summed over all six stages in trajectory 1.
green_red_log2_ratio_rep1
Diagnostic log2 ratio of green-trajectory to red-trajectory total reads in trajectory 1, with pseudocount 1.
gfp_spliced_G1_rep2
Classified spliced GFP reads in G1, trajectory replicate 2.
gfp_unspliced_G1_rep2
Classified unspliced GFP reads in G1, trajectory replicate 2.
dtom_G1_rep2
Classified dTomato reads in G1, trajectory replicate 2.
total_reads_G1_rep2
Sum of dTomato, spliced GFP, and unspliced GFP reads in G1, trajectory replicate 2.
gfp_spliced_G2_rep2
Classified spliced GFP reads in G2, trajectory replicate 2.
gfp_unspliced_G2_rep2
Classified unspliced GFP reads in G2, trajectory replicate 2.
dtom_G2_rep2
Classified dTomato reads in G2, trajectory replicate 2.
total_reads_G2_rep2
Sum of dTomato, spliced GFP, and unspliced GFP reads in G2, trajectory replicate 2.
gfp_spliced_G3_rep2
Classified spliced GFP reads in G3, trajectory replicate 2.
gfp_unspliced_G3_rep2
Classified unspliced GFP reads in G3, trajectory replicate 2.
dtom_G3_rep2
Classified dTomato reads in G3, trajectory replicate 2.
total_reads_G3_rep2
Sum of dTomato, spliced GFP, and unspliced GFP reads in G3, trajectory replicate 2.
gfp_spliced_R1_rep2
Classified spliced GFP reads in R1, trajectory replicate 2.
gfp_unspliced_R1_rep2
Classified unspliced GFP reads in R1, trajectory replicate 2.
dtom_R1_rep2
Classified dTomato reads in R1, trajectory replicate 2.
total_reads_R1_rep2
Sum of dTomato, spliced GFP, and unspliced GFP reads in R1, trajectory replicate 2.
gfp_spliced_R2_rep2
Classified spliced GFP reads in R2, trajectory replicate 2.
gfp_unspliced_R2_rep2
Classified unspliced GFP reads in R2, trajectory replicate 2.
dtom_R2_rep2
Classified dTomato reads in R2, trajectory replicate 2.
total_reads_R2_rep2
Sum of dTomato, spliced GFP, and unspliced GFP reads in R2, trajectory replicate 2.
gfp_spliced_R3_rep2
Classified spliced GFP reads in R3, trajectory replicate 2.
gfp_unspliced_R3_rep2
Classified unspliced GFP reads in R3, trajectory replicate 2.
dtom_R3_rep2
Classified dTomato reads in R3, trajectory replicate 2.
total_reads_R3_rep2
Sum of dTomato, spliced GFP, and unspliced GFP reads in R3, trajectory replicate 2.
green_stage_coverage_rep2
Number of green stages G1-G3 in trajectory 2 with at least 100 total classified reads.
red_stage_coverage_rep2
Number of red stages R1-R3 in trajectory 2 with at least 100 total classified reads.
green_total_reads_rep2
Total classified reads summed over G1-G3 in trajectory 2.
red_total_reads_rep2
Total classified reads summed over R1-R3 in trajectory 2.
trajectory_total_reads_rep2
Total classified reads summed over all six stages in trajectory 2.
green_red_log2_ratio_rep2
Diagnostic log2 ratio of green-trajectory to red-trajectory total reads in trajectory 2, with pseudocount 1.
max_stage_total_reads
Maximum total classified reads in any of the twelve FACS stages.

Quality control

Joined FACS counts to the study's trusted S4 barcode-to-intron map and applied the paper's stage-detection cutoff of at least 100 classified reads (dTomato + spliced GFP + unspliced GFP) in at least one of the twelve stage samples. This retained 6,693 elements and filtered 12,649 trusted elements below the cutoff. The paper’s stricter top-green/top-red set definitions can be reconstructed from the stage columns.

Curation notes

The FACS table retains any trusted element detectable at the paper’s ≥100-read cutoff in at least one stage, not only the final 676 top-green and 462 top-red sets. Reconstruct those final sets from stage coverage and opposite-color absence if needed. The facs_set field is an author annotation; RNA-seq count columns represent sorted bins, not direct flow-cytometry intensity values.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.