Experiment / E18LA7OSF5' UTR / Translation Efficiency MPRA (MPTA)

Integrated Toxoplasma 5′ UTR MPRA

5′ untranslated regions tune Toxoplasma translation

A 50,000-member synthetic 5′ UTR library derived from 12 endogenous Toxoplasma UTRs was cloned upstream of mNeonGreen in an RH reporter, targeted to a neutral chromosome VI locus, and sorted by mKate-positive/mNeonGreen intensity into four bins. Two biological replicates were sequenced and author QC yielded MPRA scores for 30,235 unique sequences.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

Linearized pMLP160 reporter was co-transfected with a Cas9 plasmid into RH parasites for homologous recombination at a neutral chromosome VI locus. The reporter contained the TUB1 promoter plus the first 50 nt of the TUB1 5′ UTR, the variable library sequence, an mNG-CDPK3 3′ UTR, and a constitutive mKate2-T2A-DHFR selection/gating cassette; parasites were selected with 1.5 µM pyrimethamine, sorted into mNG-negative and bottom/middle/top 15% positive bins, and scored from targeted DNA sequencing. Because high-quality RNA libraries were difficult to recover from sorted populations, RNA/DNA stability QC used bulk populations only.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (25 of 25)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 25 definitions
element_id
Unique synthetic 5′ UTR library sequence identifier, such as seq_7.
sequence
DNA sequence of the tested 5′ UTR library element as represented in the GEO/MPRA score table; the reporter also contained a constant TUB1 5′ UTR context described in the paper.
sequence_length
Length of the submitted UTR sequence in nucleotides.
uaug_count
Number of ATG (uAUG) triplets in the sequence column, i.e. the sequenced library insert.
gc_fraction
Fraction of sequence bases that are G or C.
mpra_score
Author-calculated MPRA score: the mean of a fitted normal distribution over the four FACS reporter-intensity bins, with mKate-positive/mNG-negative = 0, bottom 15% mNG-positive = 1, middle 15% = 2, and top 15% = 3.
bulk_dna_rep1_cpm
Author-provided normalized bulk parasite DNA abundance for biological replicate 1, in counts per million.
bulk_dna_rep2_cpm
Author-provided normalized bulk parasite DNA abundance for biological replicate 2, in counts per million.
bulk_dna_mean_cpm
Arithmetic mean of the two author-provided bulk DNA counts-per-million values.
bulk_dna_rep1_raw_count
Count for this sequence in the GEO counted-read archive for bulk parasite DNA replicate 1, before counts-per-million normalization.
bulk_dna_rep2_raw_count
Count for this sequence in the GEO counted-read archive for bulk parasite DNA replicate 2, before counts-per-million normalization.
mng_negative_rep1_cpm
Author-provided DNA abundance in the mKate-positive/mNG-negative FACS bin for replicate 1, normalized to matched bulk DNA counts per million.
mng_negative_rep2_cpm
Author-provided DNA abundance in the mKate-positive/mNG-negative FACS bin for replicate 2, normalized to matched bulk DNA counts per million.
mng_negative_mean_cpm
Arithmetic mean of the two normalized mKate-positive/mNG-negative bin values.
mng_bottom_rep1_cpm
Author-provided DNA abundance in the bottom 15% mNG-positive FACS bin for replicate 1, normalized to matched bulk DNA counts per million.
mng_bottom_rep2_cpm
Author-provided DNA abundance in the bottom 15% mNG-positive FACS bin for replicate 2, normalized to matched bulk DNA counts per million.
mng_bottom_mean_cpm
Arithmetic mean of the two normalized bottom-bin values.
mng_middle_rep1_cpm
Author-provided DNA abundance in the middle 15% mNG-positive FACS bin for replicate 1, normalized to matched bulk DNA counts per million.
mng_middle_rep2_cpm
Author-provided DNA abundance in the middle 15% mNG-positive FACS bin for replicate 2, normalized to matched bulk DNA counts per million.
mng_middle_mean_cpm
Arithmetic mean of the two normalized middle-bin values.
mng_top_rep1_cpm
Author-provided DNA abundance in the top 15% mNG-positive FACS bin for replicate 1, normalized to matched bulk DNA counts per million.
mng_top_rep2_cpm
Author-provided DNA abundance in the top 15% mNG-positive FACS bin for replicate 2, normalized to matched bulk DNA counts per million.
mng_top_mean_cpm
Arithmetic mean of the two normalized top-bin values.
bulk_rna_rep1_umi_count
Bulk reporter RNA UMI count for replicate 1 from the GEO counted-read archive; included as a QC/context measure and not as a sorted-bin RNA measurement.
bulk_rna_rep2_umi_count
Bulk reporter RNA UMI count for replicate 2 from the GEO counted-read archive; included as a QC/context measure and not as a sorted-bin RNA measurement.

Quality control

Reads were trimmed with cutadapt, merged with fastq-join, and unique 170–190-nt UTR sequences were retained only if detected in both unsorted bulk-DNA replicates and the input plasmid library. DESeq2 compared de-duplicated bulk RNA with bulk DNA; UTRs with adjusted p < 0.05 were excluded. Remaining UTRs required at least 25 reads in each bulk DNA replicate and at least 25 total reads across the four sorted bins in each replicate; sample counts were converted to counts per million and bin values were normalized to matched bulk counts per million. The processed table contains all 30,235 author-passing sequences; package validation confirmed A/C/G/T-only sequence, length 170–190 nt, unique IDs and sequences, finite 0–3 MPRA scores, and all raw-count thresholds.

Curation notes

This is a synthetic sequence library, not a GWAS or allele-pair experiment. The GEO score table and Supplementary Data 5 contain the same 30,235 sequence-level MPRA results. The uaug_count field is computed only from the sequence column; the reporter also had a constant TUB1 5′ UTR context described in the paper, which is not separately appended or modeled in this table. Bulk RNA UMI counts are supplied as context from the counted-read archive; MPRA bin values are author-provided counts per million normalized to corresponding bulk DNA, and mpra_score is the fitted-bin-distribution mean. No discrete genomic region or named assembly was specified for the synthetic library, so reference_genome and region_of_interest are null.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.