Experiment / E5GA41UKR5' UTR / Translation Efficiency MPRA (MPTA)

Main PLUMAGE 30-bp randomer full-length 5′-UTR variant library in PC3 and HEK293T

Multiplexed functional genomic analysis of 5′ untranslated region mutations across the spectrum of prostate cancer

The main episomal PLUMAGE library assayed 914 synthesized full-length human 5′-UTR sequences covering 545 prostate-cancer somatic mutations using a 30-bp randomer barcode linked by PacBio long reads. PC3 and HEK293T cells were measured 24 h after transfection for DNA, total mRNA, and polysome-bound mRNA; the packaged table contains the published mutation-level transcript and translation-efficiency summaries joined to unambiguous WT/mutant constructs.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated (24 h post-transfection)

Episomal pGL3-promoter luciferase reporter constructs contained full-length WT or mutant 5′-UTRs and a semi-random 30-bp barcode downstream of the luciferase coding sequence. PacBio CCS2 sequencing established barcode-to-UTR identities; Illumina PE100 sequencing quantified barcode counts in plasmid DNA, total mRNA, and polysome-bound mRNA. Transcript activity was summarized as log2(total mRNA/DNA) and translation efficiency as log2(polysome/total mRNA), with three biological replicates per cell line.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (28 of 28)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 28 definitions
variant_id
Stable variant identifier derived from the matched Supplementary Data 6a construct pair; for indels it records the sequence-catalog event.
published_effect_id
Mutant identifier from the published Supplementary Data 6d/6e effect summaries.
gene
Gene symbol for the assayed 5′-UTR.
assayed_cell_lines
Cell lines used for the main library assay; the published mutation-level summaries are aggregate and are not cell-line-stratified.
chromosome
Chromosome label from the published construct/effect annotation.
variant_position_hg19
SNV position or indel anchor position in hg19 as represented by the construct annotation.
reference_allele
Reference allele inferred from the WT/mutant construct relation and published variant label.
alternative_allele
Alternative allele inferred from the WT/mutant construct relation and published variant label.
variant_type
Sequence-level event type: SNV, deletion, or insertion.
utr_region_hg19
Broad 5′-UTR interval from the published effect annotation, in hg19 coordinates.
utr_length_bp
Length in base pairs of the matched WT full-length 5′-UTR construct.
wild_type_construct_id
Supplementary Data 6a identifier for the matched WT construct.
mutant_construct_id
Supplementary Data 6a identifier for the matched mutant construct.
wild_type_sequence
Full-length WT 5′-UTR sequence supplied in Supplementary Data 6a.
mutant_sequence
Full-length mutant 5′-UTR sequence supplied in Supplementary Data 6a.
sequence_relation
Relationship verified by direct sequence comparison between WT and mutant constructs.
sequence_difference
Number of single-base differences for SNVs, or the indel sequence-length difference for indels.
transcript_log2_fold_change
Published log2(total mRNA/DNA) mutant-versus-WT activity effect.
transcript_p_value
Published p-value for the transcript activity comparison.
transcript_fdr
Published FDR-adjusted p-value for the transcript activity comparison.
transcript_significant_fdr_lt_0_1
Whether the published transcript FDR is below 0.1.
translation_efficiency_log2_fold_change
Published log2(polysome/total mRNA) mutant-versus-WT translation-efficiency effect.
translation_efficiency_p_value
Published p-value for the translation-efficiency comparison.
translation_efficiency_fdr
Published FDR-adjusted p-value for the translation-efficiency comparison.
translation_efficiency_significant_fdr_lt_0_1
Whether the published translation-efficiency FDR is below 0.1.
functional_in_any_layer
Whether the published FDR is below 0.1 in either transcript activity or translation efficiency.
qc_pass
Package-level QC flag; only true rows are present in the processed table.
source_data
Supplementary Data 6a, 6d, and 6e and the deposited GEO series used to construct or audit the row.

Quality control

The paper required high-quality PacBio CCS2 reads (minimum three passes and predicted accuracy 0.9), exact matching of consensus 5′-UTRs to expected constructs, barcode extraction with matching 4-nt flanks, and at least 0.5 CPM for barcodes used in ratio calculations. It reports more than 96% of extracted barcodes matching the PacBio catalog and replicate correlations above 0.99 for DNA, 0.8 for total RNA, and 0.89 for polysome RNA. For this package, only published mutation-effect rows that could be joined unambiguously to one WT and one mutant sequence in Supplementary Data 6a were retained: 536 rows passed, while the duplicate CHRM2_C_G..._2 comparison and the DCAF6_T_G comparison with a non-clean sequence relation were excluded. Nonsignificant published rows remain in the table; FDR < 0.1 is represented by explicit Boolean columns.

Curation notes

This is one aggregate main-library experiment because the published Supplementary Data 6d/6e effect tables do not provide separate PC3 and HEK293T effect columns; biosample_id is therefore null rather than assigning aggregate statistics to one cell line. Raw GEO GSE149487 includes separate PC3 and HEK293T DNA, total-RNA, and polysome count files, but the deposited package does not provide a sufficiently direct cell-line-specific reanalysis table for these published mutation-level effects. The seven special 8-bp rows in Supplementary Data 6d/6e are represented in the separate pilot experiment where applicable. The MAT1A construct is retained as an insertion based on its actual sequence comparison despite the source identifier's deletion-like label; its source annotation is preserved in the variant_id and source fields.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.