Experiment / E0OVC1NQIOther

Synthetic 3′UTR GFP MPRA in primary CD8+ T cells

Learning the sequence code of protein expression in human immune cells

A 467-member synthetic 3′UTR library, with six repeats of each designed motif in a GFP reporter, was delivered by an integrating retroviral vector to primary CD8+ T cells from three healthy donors. The sorted GFP-high, GFP-low, and total GFP-positive fractions were quantified by genomic-DNA amplicon sequencing.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

αCD3/αCD28 activation for 48h before retroviral transduction; IL-2/IL-15/IL-7 expansion

Primary CD8+ T cells from three healthy donors were activated with αCD3/αCD28 for 48h before retroviral transduction and expanded with IL-2, IL-15, and IL-7. The library contained 467 synthetic 200-nt 3′UTR oligos: six repeats of each designed motif, including two 7-nt validation motifs, with fixed cloning flanks. Transduction was maintained at 8–12% to minimize multiple integrations; the top and bottom 15% GFP-expressing cells and the full GFP-positive fraction were sorted. Read2 was trimmed by removing 39 5′ bases and 57 3′ bases, aligned to the UTR library with Bowtie2 local mode (--very-sensitive), and counted with htseq-count. The author-normalized source counts are scaled to 10,000 reads per library; high-versus-low log2 enrichment columns in this package are derived from those normalized values.

Retroviral, genomically integrating GFP–synthetic 3′UTR reporter library (pRETRO-SUPER_GFP; not lentiviral); cells were transduced, GFP-high/GFP-low populations were FACS-sorted, and integrated 3′UTR abundance was quantified from genomic-DNA amplicon sequencing.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (49 of 49)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 49 definitions
element_id
Unique source oligo identifier for the tested synthetic 3′UTR element.
motif_sequence
Designed motif sequence encoded in the element identifier.
motif_length_nt
Length of the designed motif in nucleotides (6 or 7).
library_design
Synthetic library design class: six-repeat 6-mer motif or six-repeat 7-mer validation motif.
reporter_oligo_sequence
Complete 200-nt synthetic reporter oligo sequence from supplementary Data S4, including fixed flanking cloning sequence.
designed_repeat_count
Number of intended motif repeats in the synthetic oligo (six).
motif_occurrence_count_in_full_oligo
Literal non-overlapping occurrences of the motif in the full 200-nt oligo; this can exceed the designed repeat count when the fixed sequence contains the motif.
a_count
Number of adenines in the designed motif.
c_count
Number of cytosines in the designed motif.
g_count
Number of guanines in the designed motif.
t_count
Number of thymines in the designed motif.
ag_count
A+G count in the designed motif.
ct_count
C+T count in the designed motif.
gc_count
G+C count in the designed motif.
ag_fraction
AG count divided by motif length.
ct_fraction
CT count divided by motif length.
gc_fraction
GC count divided by motif length.
raw_count_all_27_libraries
Sum of raw read counts for this element across all 27 deposited samples; used for the 40-read source QC threshold.
qc_pass
Whether the element passed the package QC filters and was retained in the table.
gfp_high_rep1_raw_count
Raw genomic-DNA amplicon read count in the GFP-high fraction for replicate 1 (HEK batch 1 or T-cell donor 1).
gfp_high_rep2_raw_count
Raw genomic-DNA amplicon read count in the GFP-high fraction for replicate 2 (HEK batch 2 or T-cell donor 2).
gfp_high_rep3_raw_count
Raw genomic-DNA amplicon read count in the GFP-high fraction for replicate 3 (HEK batch 3 or T-cell donor 3).
gfp_low_rep1_raw_count
Raw genomic-DNA amplicon read count in the GFP-low fraction for replicate 1 (HEK batch 1 or T-cell donor 1).
gfp_low_rep2_raw_count
Raw genomic-DNA amplicon read count in the GFP-low fraction for replicate 2 (HEK batch 2 or T-cell donor 2).
gfp_low_rep3_raw_count
Raw genomic-DNA amplicon read count in the GFP-low fraction for replicate 3 (HEK batch 3 or T-cell donor 3).
gfp_total_rep1_raw_count
Raw genomic-DNA amplicon read count in the full GFP-positive fraction for replicate 1 (HEK batch 1 or T-cell donor 1).
gfp_total_rep2_raw_count
Raw genomic-DNA amplicon read count in the full GFP-positive fraction for replicate 2 (HEK batch 2 or T-cell donor 2).
gfp_total_rep3_raw_count
Raw genomic-DNA amplicon read count in the full GFP-positive fraction for replicate 3 (HEK batch 3 or T-cell donor 3).
gfp_high_rep1_normalized_count_10k
Author-normalized GFP-high read count for replicate 1, scaled to 10,000 library reads.
gfp_high_rep2_normalized_count_10k
Author-normalized GFP-high read count for replicate 2, scaled to 10,000 library reads.
gfp_high_rep3_normalized_count_10k
Author-normalized GFP-high read count for replicate 3, scaled to 10,000 library reads.
gfp_low_rep1_normalized_count_10k
Author-normalized GFP-low read count for replicate 1, scaled to 10,000 library reads.
gfp_low_rep2_normalized_count_10k
Author-normalized GFP-low read count for replicate 2, scaled to 10,000 library reads.
gfp_low_rep3_normalized_count_10k
Author-normalized GFP-low read count for replicate 3, scaled to 10,000 library reads.
gfp_total_rep1_normalized_count_10k
Author-normalized full GFP-positive read count for replicate 1, scaled to 10,000 library reads.
gfp_total_rep2_normalized_count_10k
Author-normalized full GFP-positive read count for replicate 2, scaled to 10,000 library reads.
gfp_total_rep3_normalized_count_10k
Author-normalized full GFP-positive read count for replicate 3, scaled to 10,000 library reads.
gfp_high_mean_normalized_count_10k
Arithmetic mean of the three author-normalized GFP-high counts.
gfp_high_sd_normalized_count_10k
Sample standard deviation of the three author-normalized GFP-high counts.
gfp_low_mean_normalized_count_10k
Arithmetic mean of the three author-normalized GFP-low counts.
gfp_low_sd_normalized_count_10k
Sample standard deviation of the three author-normalized GFP-low counts.
gfp_total_mean_normalized_count_10k
Arithmetic mean of the three author-normalized full GFP-positive counts.
gfp_total_sd_normalized_count_10k
Sample standard deviation of the three author-normalized full GFP-positive counts.
log2_enrichment_high_vs_low_rep1
Derived log2(GFP-high normalized count / GFP-low normalized count) for replicate 1.
log2_enrichment_high_vs_low_rep2
Derived log2(GFP-high normalized count / GFP-low normalized count) for replicate 2.
log2_enrichment_high_vs_low_rep3
Derived log2(GFP-high normalized count / GFP-low normalized count) for replicate 3.
log2_enrichment_high_vs_low_mean
Arithmetic mean of the three replicate-specific high-versus-low log2 enrichments.
log2_enrichment_high_vs_low_sd
Sample standard deviation of the three replicate-specific high-versus-low log2 enrichments.
log2_enrichment_high_vs_low_se
Standard error of the mean high-versus-low log2 enrichment, calculated as sample SD divided by sqrt(3).

Quality control

Author QC trimmed read2, aligned with Bowtie2 --very-sensitive, counted with htseq-count, excluded motifs with fewer than 40 reads across all samples, and normalized each library to 10,000 reads. Package QC additionally required a unique element ID, a matching 200-nt oligo sequence, complete high/low/total counts for all three replicates, finite positive normalized values, and at least 40 raw reads summed across all 27 deposited libraries. All 467/467 library elements passed (minimum all-library total 315); no rows were removed.

Curation notes

Three primary CD8+ T-cell donors are represented as replicate 1–3 (GEO high/low/total triplets: GSM8840945/GSM8840954/GSM8840955, GSM8840956/GSM8840957/GSM8840958, and GSM8840959/GSM8840960/GSM8840961). This package treats the shared library in this biosample as one child experiment so each experiment table has one cell context. The high and low columns are FACS fractions rather than RNA/DNA ratios: high and low are the top and bottom 15% of GFP-positive cells; total is the full GFP-positive fraction. Normalized counts are the author/GEO 10,000-read values and include a pseudocount (source raw zeros are positive in the normalized table). log2_enrichment_high_vs_low_* is derived per replicate from these normalized counts. The full oligo sequence is 200 nt and contains six designed motif repeats; some motifs occur an additional time in fixed flanking sequence, so both designed_repeat_count and motif_occurrence_count_in_full_oligo are retained. Source rows from supplementary Data S4 and GEO matched exactly for all 467 elements and 27 samples. The assay is classified as Other because the paper specifies a genomically integrating retroviral vector, not lentivirus.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.