Experiment / E4OT11CO5Sort-Seq / Flow-Seq MPRA

C3U Flp-In-293 pooled 3′-UTR activity screen

Systematic identification of regulatory elements in conserved 3′-untranslated regions of human transcripts

Two independently generated C3U libraries of 16,332 conserved human 34-nt 3′-UTR elements were tested in pooled Flp-In-293 cells. FACS expression bins and high-throughput sequencing of reporter-insert amplicons were used to identify elements enriched in high- or low-DIR populations.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

A bidirectional CMV-GFP/mCherry reporter carrying each insert in the mCherry 3′ UTR was recombinase-integrated at the single FRT locus of Flp-In-293 cells. GFP-positive cells were sorted into approximately 10% DIR bins (H10, H20, H30, H40, L10, L20, L30, L40), and genomic-DNA amplicons were sequenced; GEO supplies technical-replicate-consolidated and adjacent-bin-merged counts and log2 frequencies for independent library constructions A and B.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (42 of 42)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 42 definitions
element_id
Unique C3U library sequence identifier.
sequence
The tested 34-nt conserved 3′-UTR sequence.
sequence_length
Length of the tested sequence in nucleotides.
published_prediction_id
Publisher Table S2 ranked prediction identifier, such as C3U-R1 or C3U-A1; blank when not in Table S2.
published_activity_call
Publisher Table S2 call: repressor or activator.
published_prediction_rank
Numeric rank parsed from the publisher prediction identifier.
published_log_high_vs_low
Publisher Table S2 log2(high-DIR frequency/low-DIR frequency) value.
published_q_value
Publisher Table S2 Benjamini-Hochberg q-value.
library_a_background_count
Library A background-population count supplied by GEO.
library_b_background_count
Library B background-population count supplied by GEO.
library_a_h10_h20_count
Library A raw read count in the merged H10/H20 high-DIR bins.
library_a_h10_h20_log2_relative_freq
Library A GEO-supplied log2 normalized frequency relative to background in H10/H20.
library_a_h30_h40_count
Library A raw read count in the merged H30/H40 high-DIR bins.
library_a_h30_h40_log2_relative_freq
Library A GEO-supplied log2 normalized frequency relative to background in H30/H40.
library_a_l10_l20_count
Library A raw read count in the merged L10/L20 low-DIR bins.
library_a_l10_l20_log2_relative_freq
Library A GEO-supplied log2 normalized frequency relative to background in L10/L20.
library_a_l30_l40_count
Library A raw read count in the merged L30/L40 low-DIR bins.
library_a_l30_l40_log2_relative_freq
Library A GEO-supplied log2 normalized frequency relative to background in L30/L40.
library_b_h10_h20_count
Library B raw read count in the merged H10/H20 high-DIR bins.
library_b_h10_h20_log2_relative_freq
Library B GEO-supplied log2 normalized frequency relative to background in H10/H20.
library_b_h30_h40_count
Library B raw read count in the merged H30/H40 high-DIR bins.
library_b_h30_h40_log2_relative_freq
Library B GEO-supplied log2 normalized frequency relative to background in H30/H40.
library_b_l10_l20_count
Library B raw read count in the merged L10/L20 low-DIR bins.
library_b_l10_l20_log2_relative_freq
Library B GEO-supplied log2 normalized frequency relative to background in L10/L20.
library_b_l30_l40_count
Library B raw read count in the merged L30/L40 low-DIR bins.
library_b_l30_l40_log2_relative_freq
Library B GEO-supplied log2 normalized frequency relative to background in L30/L40.
library_a_high_mean_log2_relative_freq
Mean Library A log2 relative frequency across high-DIR bins with >20 reads.
library_a_low_mean_log2_relative_freq
Mean Library A log2 relative frequency across low-DIR bins with >20 reads.
library_a_high_vs_low_log2_effect
Library A high-DIR mean minus low-DIR mean; positive indicates high-DIR enrichment.
library_b_high_mean_log2_relative_freq
Mean Library B log2 relative frequency across high-DIR bins with >20 reads.
library_b_low_mean_log2_relative_freq
Mean Library B log2 relative frequency across low-DIR bins with >20 reads.
library_b_high_vs_low_log2_effect
Library B high-DIR mean minus low-DIR mean; positive indicates high-DIR enrichment.
combined_high_mean_log2_relative_freq
Mean of all available high-DIR log2 relative frequencies from bins with >20 reads.
combined_low_mean_log2_relative_freq
Mean of all available low-DIR log2 relative frequencies from bins with >20 reads.
combined_high_vs_low_log2_effect
Combined high-DIR mean minus low-DIR mean; positive indicates high-DIR enrichment and negative indicates low-DIR enrichment.
replicate_effect_difference
Library A high-vs-low effect minus Library B high-vs-low effect when both library-specific effects are estimable.
n_high_bins_gt20
Number of high-DIR merged-bin files with more than 20 reads across libraries A and B.
n_low_bins_gt20
Number of low-DIR merged-bin files with more than 20 reads across libraries A and B.
n_library_replicates_with_both_sides
Number of library constructions for which both a high-DIR and a low-DIR bin exceeded 20 reads.
effect_direction
Direction from the combined derived effect: high_DIR_enriched, low_DIR_enriched, or neutral.
published_direction_concordance
Whether the derived effect direction agrees with the publisher's repressor/activator call.
qc_pass
TRUE for rows retained after the package QC filter.

Quality control

The authors retained reads containing both terminal universal adapters, matched reads to the synthetic C3U library, consolidated technical replicates, merged adjacent bins, normalized each sorted-bin frequency to its library background population, required more than 20 counts in a sample for downstream sequence analysis, and excluded sequences seen only in background. They called repressors or activators with a high-versus-low DIR frequency difference of at least twofold and Benjamini-Hochberg q-value <0.05. For this package, a row was retained only if its sequence was 34 nt of A/C/G/T, had nonzero background representation in at least one library, and had >20 reads in at least one high-DIR and one low-DIR merged-bin file across the two libraries; summary effects use only bins passing the >20-count threshold. This retained 5,668 of the 12,970 sequence IDs present in the GEO processed files.

Curation notes

The paper presents a custom pooled post-transcriptional reporter screen rather than using the MPRA acronym, but it meets the MPRA definition operationally: thousands of regulatory inserts are assayed in parallel by a reporter, sorted by expression, and quantified by sequencing. The primary classification is Sort-Seq / Flow-Seq MPRA because FACS binning is the assay readout; the single-locus Flp recombinase integration is documented in assay_type_additional_details. Libraries A and B are independent library constructions under the same HEK293 condition and are packaged as one experiment. GEO's published log2 relative frequencies are preserved for every bin, while derived means ignore bins with <=20 reads. The combined effect can be estimable even when no individual library has both sides covered; those cases are marked by n_library_replicates_with_both_sides and have blank library-specific effect differences where appropriate. Table S2 annotations are joined by element_id when available.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.