Experiment / E4SOLQBSNEpisomal Plasmid MPRA

PolyU+/− random intron validation MPRA

Sequence-dependent and -independent effects of intron-mediated enhancement learned from thousands of random introns

A smaller validation library tested 30 parent random introns in original, U4-deleted, and U4-added sequence forms, for 90 designed intron sequences, together with intronless controls. The pool was transiently transfected into HEK293T A2 cells and assayed by two-replicate RNA-seq of dTomato, spliced GFP, and unspliced GFP reporter reads.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Transient transfection; polyU/U4 deletion or addition variants

This validation library was the exception to the study’s integrated reporter experiments: the pool was transiently transfected, and RNA was collected for two replicate reporter-count measurements. Each oligo uses a dual dTomato/GFP reporter with a barcode; the 30 parent introns were minimally permuted to disperse or cluster U residues, producing the original, polyU4-deleted, and polyU4-added conditions. polyU counts in this package count overlapping TTT/TTTT/TTTTT motifs in the variable intron region.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (31 of 31)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 31 definitions
element_id
Stable package identifier for the oligo-level validation element.
oligo_id
Source oligo identifier from the GEO workbook.
description
Source library description identifying the parent intron or intronless control.
barcode
Reporter barcode associated with the oligo.
parent_intron
Normalized parent random-intron identifier for variants; blank for intronless controls.
condition
Normalized design condition: original, polyU4_deleted, polyU4_added, or intronless_control.
intron_sequence
Full deposited intron sequence; blank for intronless controls.
random_region_sequence
Internal variable sequence after removing the constant splice-site flanks; blank for controls.
intron_length
Length of the full deposited intron sequence in nucleotides; blank for controls.
random_region_length
Length of the variable internal sequence in nucleotides; blank for controls.
polyU3_count
Number of overlapping TTT motifs in the variable region.
polyU4_count
Number of overlapping TTTT motifs in the variable region.
polyU5_count
Number of overlapping TTTTT motifs in the variable region.
random_region_gc_fraction
Fraction of variable-region bases that are G or C.
count_completeness
Whether both replicate spliced-GFP source cells were populated (complete) or at least one was blank (missing_spliced_GFP).
dtom_rep1
Deposited dTomato read count for replicate 1.
gfp_spliced_rep1
Deposited spliced-GFP read count for replicate 1; blank source cells are preserved as blank.
gfp_unspliced_rep1
Deposited unspliced-GFP read count for replicate 1.
gfp_total_rep1
Spliced plus unspliced GFP count for replicate 1; a blank source spliced-GFP cell contributes zero only to this derived total.
total_reads_rep1
dTomato plus total GFP classified reads for replicate 1.
raw_log2_gfp_dtom_ratio_rep1
Non-normalized log2((1 + total GFP) / (1 + dTomato)) for replicate 1; blank when spliced-GFP data are incomplete.
splicing_efficiency_rep1
Spliced-GFP divided by total GFP for replicate 1; blank when spliced-GFP data are incomplete or total GFP is zero.
dtom_rep2
Deposited dTomato read count for replicate 2.
gfp_spliced_rep2
Deposited spliced-GFP read count for replicate 2; blank source cells are preserved as blank.
gfp_unspliced_rep2
Deposited unspliced-GFP read count for replicate 2.
gfp_total_rep2
Spliced plus unspliced GFP count for replicate 2; a blank source spliced-GFP cell contributes zero only to this derived total.
total_reads_rep2
dTomato plus total GFP classified reads for replicate 2.
raw_log2_gfp_dtom_ratio_rep2
Non-normalized log2((1 + total GFP) / (1 + dTomato)) for replicate 2; blank when spliced-GFP data are incomplete.
splicing_efficiency_rep2
Spliced-GFP divided by total GFP for replicate 2; blank when spliced-GFP data are incomplete or total GFP is zero.
mean_raw_log2_gfp_dtom_ratio
Mean of the two replicate raw ratios; blank when either replicate has incomplete spliced-GFP data.
min_total_reads_across_replicates
Smaller of the two replicate total classified-read counts.

Quality control

Applied the study's stated polyU-library filter of at least 100 total classified reads per replicate, where total is dTomato + spliced GFP + unspliced GFP. This retained 278 of 280 source rows and excluded 2 low-count rows (o_42408, o_42806). Source-cell blanks for spliced GFP are preserved and are not imputed.

Curation notes

The direct GEO workbook contains 280 rows: 90 original parent introns, 90 U4-deleted designs, 90 U4-added designs, and 10 intronless controls. For U4-added/deleted rows, the deposited spliced-GFP cells are blank while low unspliced-GFP counts are present; this package preserves those blanks and does not substitute unavailable values. The public supplementary S9 sheet has one fewer row than the direct GEO workbook, so the direct GEO workbook is the source for this table.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.