Experiment / E561K447IEpisomal Plasmid MPRA

HEK293T episomal MPRA of rare noncoding variant alleles

Massively parallel identification of functionally consequential noncoding genetic variants in undiagnosed rare disease patients

An episomal MPRA library of 3,059 rare noncoding single-nucleotide variants was tested in human HEK293T cells. Each variant was represented by four 100-nucleotide allele sequences (reference, patient variant, and two alternative alleles), with five 12-nucleotide barcodes per construct and two plasmid-DNA plus two cDNA sequencing replicates.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

The authors synthesized 100-nt genomic sequences centered on each rare noncoding SNV and generated four alleles at every site: the reference base, the patient variant base, and the two other possible nucleotides. Each sequence was linked to five independent 12-nt barcodes. The oligos were cloned upstream of an SV40 promoter-GFP reporter in a plasmid library, transfected into HEK293T cells, and harvested 24 h later; targeted sequencing of barcode-containing reporter transcripts measured cDNA abundance. MPRA plasmid DNA was sequenced in two technical replicates and expressed cDNA in two biological replicates, and DESeq2 compared cDNA with the plasmid pool.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (42 of 42)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 42 definitions
element_id
Stable package identifier for one allele construct, assigned in GEO source-matrix row order.
source_sequence_id
Original GEO/Supplementary Table S2 sequenceID for the allele construct.
variant_id
Source variant identifier in the form chromosome:1-based-position-reference-variant-alleles.
chromosome
Chromosome component of the source variant identifier.
position_grch37_1based
1-based variant position in the GRCh37 coordinate system used by the source annotations.
reference_allele
Reference nucleotide encoded by the final two-letter allele suffix of variant_id.
variant_allele
Patient variant nucleotide encoded by the final two-letter allele suffix of variant_id and the S1 Variant construct.
tested_allele
Nucleotide represented by this construct at the variant position.
allele_role
Source S1 oligoType: Reference, Variant, or Alternative.
variant_class
Source S1 variant classification: DeNovo or Inherited.
proband
Source patient identifier associated with the candidate variant.
genomic_location
Source genomic annotation for the candidate variant, such as intergenic, intronic, or UTR.
gene
Source gene annotation; NA indicates that no gene value was supplied.
nearest_candidate_gene_tss
Nearest candidate gene/TSS annotation from Supplementary Table S1.
distance_to_candidate_gene_tss
Source distance in base pairs to the nearest candidate gene TSS.
gnomad_maf
gnomAD minor allele frequency from Supplementary Table S1; NA indicates unavailable.
internal_maf
Internal cohort allele frequency from Supplementary Table S1.
cadd
Source CADD score; NA indicates unavailable.
cadd_phred
Source CADD PHRED-scaled score; NA indicates unavailable.
gerp
Source GERP conservation score; NA indicates unavailable.
genomic_sequence_100bp
Uppercase 100-nt genomic sequence context from S1, with the tested allele at position 50 (0-based).
variant_position_in_sequence_0based
0-based position of the tested allele in genomic_sequence_100bp; 50 for all constructs.
n_barcodes
Number of unique S1 oligo barcodes assigned to the construct; 5 for every retained construct.
pDNA_rep1_normalized_count
GEO MPRA plasmid-DNA normalized count for technical replicate 1.
pDNA_rep2_normalized_count
GEO MPRA plasmid-DNA normalized count for technical replicate 2.
cDNA_rep1_normalized_count
GEO MPRA expressed cDNA normalized count for biological replicate 1.
cDNA_rep2_normalized_count
GEO MPRA expressed cDNA normalized count for biological replicate 2.
base_mean
Supplementary Table S2 DESeq2 baseMean across the two pDNA and two cDNA samples.
activity_log2_fold_change
Supplementary Table S2 DESeq2 log2 fold change for cDNA versus pDNA activity.
activity_lfc_se
Supplementary Table S2 standard error of the activity log2 fold change.
activity_stat
Supplementary Table S2 DESeq2 test statistic for cDNA versus pDNA activity.
activity_p_value
Supplementary Table S2 nominal p-value for cDNA versus pDNA activity.
activity_p_adj
Supplementary Table S2 multiple-testing-adjusted p-value for cDNA versus pDNA activity.
activity_significant_padj_0_05
Boolean equivalent of the source S2 Sig field: true when activity_p_adj is below 0.05.
activity_neg_log10_p_adj
Supplementary Table S2 negative base-10 logarithm of the adjusted activity p-value.
variant_reference_activity_log2_fold_change
S3 activity log2 fold change for the reference construct, repeated for all four constructs of the variant.
variant_allele_activity_log2_fold_change
S3 activity log2 fold change for the patient variant construct, repeated for all four constructs of the variant.
variant_effect_log2_fold_change
S3 log2 fold-change difference of patient variant activity minus reference activity.
variant_effect_z_score
S3 z-score for the patient-variant versus reference activity effect.
variant_effect_significant_abs_z_2
Boolean indicating whether the S3 variant-effect absolute z-score is greater than 2.
variant_effect_direction
Derived direction from the S3 z-score: increased, decreased, or not_significant.
source_tables
Row-level provenance: GEO MPRA matrix and Supplementary Tables S1-S3.

Quality control

The authors summed reads for five barcodes per sequence and used DESeq2 to compare cDNA abundance with the plasmid-DNA pool. The source Supplementary Table S2 marks construct activity at adjusted p-value < 0.05, and Supplementary Table S3 defines variant effects using an absolute z-score > 2. Package QC required a unique construct identifier, an exact join across the GEO MPRA matrix and Supplementary Tables S1-S3, exactly five unique barcodes per construct, a 100-nt A/C/G/T genomic sequence with the tested allele at position 50 (0-based), and finite positive normalized counts plus finite DESeq2 and variant-effect statistics. All 12,236 constructs passed these checks and no rows were removed; the package retains non-significant constructs for complete variant-level context.

Curation notes

This is the study's single MPRA experiment; the ancillary bulk RNA-seq/CRISPRi validation is preserved in raw_data but is not represented as a separate MPRA experiment. The processed table has one row per allele construct (12,236 rows: 3,059 variants x 4 alleles) and integrates the GEO normalized pDNA/cDNA counts with S1 design annotations, S2 construct-level DESeq2 results, and S3 variant-level effects. The 138 variants with absolute z-score > 2 comprise 91 de novo and 47 inherited variants; the 91 de novo effects split into 51 decreased and 40 increased reporter activity, matching the paper's reported MPRA result. The paper states that variants were called on GRCh38 and lifted over to GRCh37 for genomic analyses; the source variant IDs use those GRCh37-lifted coordinates, while the supplied 100-nt sequences are retained as provided. GRCh37 is used here for the coordinate metadata. NA is used in the CSV for missing source annotations.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.