Experiment / E7AS78FO0Targeted Genomic Integration MPRA

Recombinase-mediated 5′ UTR screen in HEK 293T landing-pad cells

High-throughput 5′ UTR engineering for enhanced protein production in non-viral gene therapies

A 12,000-member library of 100-bp natural and computationally designed 5′ UTRs was cloned upstream of a GFP reporter and integrated as single copies into Bxb1 landing-pad HEK 293T cells. Two independently screened landing-pad cell lines were sorted into GFP-expression bins, and genomic amplicon sequencing with DESeq2 quantified UTR enrichment in the top 0–2.5%, 2.5–5%, and 5–10% bins relative to unsorted cells.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Bxb1 recombinase co-transfection with puromycin selection; GFP-based FACS sorting

The reporter plasmid carried an attB site, the 5′ UTR library upstream of a GFP coding sequence with a strong Kozak context, and an RFP/puromycin selection cassette. Bxb1-mediated attB × attP recombination placed one payload copy at a defined genomic landing pad in HEK-LP3 and HEK-LP9 derivatives. Cells were maintained at greater than 25-fold library coverage, sorted into three high-GFP brackets plus an unsorted 0–100% control, and genomic DNA amplicons were sequenced on an Illumina NextSeq. The released processed files are DESeq2 comparisons of each sorted bin against the unsorted control; the table joins all three comparisons to the 140-bp oligo sequence and design annotations.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (38 of 38)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 38 definitions
utr_id
Library UTR identifier; the source workbooks incorrectly label this field ensembl_gene_id.
library_class
Whether the library member is naturally occurring, a two-generation synthetic test sequence, or a genetic-algorithm-optimized synthetic sequence.
sequence_140bp
Complete synthesized oligo sequence, including 20-bp flanking regions on each side of the 100-bp variable UTR.
utr_sequence_100bp
The variable 100-bp 5′ UTR sequence extracted from sequence_140bp.
natural_source_transcript
Transcript or source identifier reported for naturally occurring library members.
natural_source_cell_type_reported
Cell/tissue or other source value reported in Supplementary Data 1; some rows contain source transcript-like values because the publisher annotation is inconsistent.
natural_source_rnaseq_rpkm
RNA-seq RPKM used in the natural-sequence selection, when supplied.
natural_source_translation_efficiency
Translation efficiency (Ribo-seq RPKM divided by RNA-seq RPKM) supplied for the natural sequence, when supplied.
natural_source_gene_name
Gene name supplied for the natural sequence, when supplied.
natural_source_strand
Strand value supplied for the natural sequence.
synthetic_design_score
Model-predicted score reported for a synthetic sequence.
synthetic_design_model
Cell/source and predicted target represented by the synthetic design model, such as HEK_Andrev2015_TE or pc3_Ribo.
synthetic_design_generation
Genetic-algorithm generation at which the sequence was selected.
synthetic_design_information
Publisher label for the synthetic design selection, such as bestTop, Up, or Down.
synthetic_input_sequence_100bp
The 100-bp starting sequence for the two-generation synthetic test set, when supplied.
synthetic_input_score
Model score of the starting sequence for a synthetic design, when supplied.
bin_0_2_5_base_mean
DESeq2 baseMean for the 0–2.5% GFP bin versus the unsorted 0–100% control.
bin_0_2_5_log2_fold_change
DESeq2 log2 fold change for UTR enrichment in the 0–2.5% GFP bin versus unsorted control.
bin_0_2_5_lfc_se
Standard error of the 0–2.5% bin log2 fold change.
bin_0_2_5_stat
DESeq2 test statistic for the 0–2.5% bin comparison.
bin_0_2_5_p_value
Raw DESeq2 p-value for the 0–2.5% bin comparison.
bin_0_2_5_p_adj
DESeq2 multiple-testing-adjusted p-value for the 0–2.5% bin comparison.
bin_2_5_5_base_mean
DESeq2 baseMean for the 2.5–5% GFP bin versus the unsorted 0–100% control.
bin_2_5_5_log2_fold_change
DESeq2 log2 fold change for UTR enrichment in the 2.5–5% GFP bin versus unsorted control.
bin_2_5_5_lfc_se
Standard error of the 2.5–5% bin log2 fold change.
bin_2_5_5_stat
DESeq2 test statistic for the 2.5–5% bin comparison.
bin_2_5_5_p_value
Raw DESeq2 p-value for the 2.5–5% bin comparison.
bin_2_5_5_p_adj
DESeq2 multiple-testing-adjusted p-value for the 2.5–5% bin comparison.
bin_5_10_base_mean
DESeq2 baseMean for the 5–10% GFP bin versus the unsorted 0–100% control.
bin_5_10_log2_fold_change
DESeq2 log2 fold change for UTR enrichment in the 5–10% GFP bin versus unsorted control.
bin_5_10_lfc_se
Standard error of the 5–10% bin log2 fold change.
bin_5_10_stat
DESeq2 test statistic for the 5–10% bin comparison.
bin_5_10_p_value
Raw DESeq2 p-value for the 5–10% bin comparison.
bin_5_10_p_adj
DESeq2 multiple-testing-adjusted p-value for the 5–10% bin comparison.
mean_log2_enrichment
Mean of the three released bin log2 fold changes.
minimum_log2_enrichment
Smallest of the three released bin log2 fold changes.
paper_validation_candidate
Boolean flag for the 13 UTR IDs listed in Supplementary Table 2 as selected for downstream validation.
validated_utr_name
Named validation construct for the three lead UTRs: NeoUTR1, NeoUTR2, or NeoUTR3; blank otherwise.

Quality control

The authors inspected FASTQ files with FastQC, filtered and trimmed reads with fastx_clipper, collapsed identical reads with fastx_collapser, aligned with Bowtie2 in very-sensitive mode, retained mapped reads with SAMtools, normalized sample counts using DESeq2 size factors, assessed replicate reproducibility with Pearson correlation, and identified differential UTR enrichment with DESeq2. Package QC retained 11,651 UTRs: the ID had to be present in the complete 12,000-member library and all three released bin workbooks, the 140-bp library sequence had to be present, and all baseMean, log2FoldChange, lfcSE, stat, pvalue, and padj values had to be finite with baseMean > 0 and pvalue/padj in [0,1]. This excludes 349 library IDs absent from the released workbooks and the orphan analysis ID 12001; it does not exclude non-significant but valid UTRs.

Curation notes

This package represents the recombinase/landing-pad screen because the three released GEO binned workbooks are the processed DESeq2 results used for the paper’s high-GFP candidate selection; the filenames themselves do not identify Rec versus Lent. GEO also contains eight Lent and eight Rec raw SRA samples, but the lentiviral screen has no separately released processed table and raw FASTQ/SRA data were intentionally not copied into the package. The source binned header ensembl_gene_id is a misnomer: values are library UTR numbers, including an analysis-only 12001 that has no corresponding sequence in the 12,000-member library. The paper’s text reports adjusted-p-value selection of 13 candidates, while the released bin workbooks have no padj below 0.05; the exact 13 IDs from Supplementary Table 2 are therefore retained as paper_validation_candidate rather than recomputed from the inconsistent padj fields. HEK 293T landing-pad clones are not separately registered in Cellosaurus, so the parental HEK293T accession CVCL:0063 is used.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.