Experiment / E11UT7DTOTargeted Genomic Integration MPRA

COMT protein abundance flow-cytometry MAVE

Integrated multiplexed assays of variant effect reveal determinants of catechol-O-methyltransferase gene expression

A codon-saturation COMT library covering MB-COMT codons 40–74 and 136–158 was stably integrated into the HEK293T landing-pad cell line. Four biological replicates were sorted by flow cytometry and profiled by amplicon sequencing; this table reports the published ALDEx2 effects for the low-protein flow population P3 relative to total genomic DNA.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Doxycycline induction (2 µg/mL for 21 h) after AP1903 selection; fresh-media stimulation for 1 h 15 min before flow cytometry

This is a targeted genomic-integration saturation-mutagenesis reporter assay rather than a barcode-based episomal MPRA. Bxb1 recombinase inserted the pDEST_HC_Rec_Bxb_v2_UTR5_Flag_COMT_moxGFP_LPS construct into a landing pad, with moxGFP reporting protein abundance and IRES-mCherry serving as a transcript/induction control; all silent and missense substitutions and amber nonsense substitutions were designed across two COMT coding ROIs. The selected flow readout is population P3 relative to total gDNA; population P4 was excluded in the source analysis. MAVEdb defines positive protein-abundance scores as lower protein abundance. Variant effects are published ALDEx2 standardized differences with 75% confidence intervals and FDR.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (38 of 38)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 38 definitions
mavedb_accession
MAVEdb score-set accession and variant row identifier
hgvs_nt
Target-relative nucleotide HGVS from MAVEdb
hgvs_pro_mavedb
Target-relative protein HGVS from MAVEdb; the target begins at MB-COMT Glu40
variant_id
Unique plasmid-reference variant identifier from Dataset EV3
roi
Mutagenesis region: ROI_1 is codons 40–74 and ROI_2 is codons 136–158
position_plasmid
1-based plasmid reference coordinate of the first altered base
ref_nt
Reference nucleotide or multi-nucleotide sequence
alt_nt
Alternate nucleotide or multi-nucleotide sequence
ref_codon
Reference codon
alt_codon
Alternate codon
aa_position
MB-COMT amino-acid position after the N-terminal tag adjustment
ref_aa
Reference amino acid in one-letter code
alt_aa
Alternate amino acid in one-letter code; * denotes stop
aa_change
MB-COMT protein consequence, for example p.E40K
variant_type
Silent, Missense, or Nonsense consequence
mutation_type
SNP, di_nt_MNP, or tri_nt_MNP substitution class
polysome_fraction
NA because this is the protein-abundance experiment
effect_score
Author Dataset EV3 full-precision ALDEx2 standardized effect for flow P3 versus total gDNA; cross-checked to the MAVEdb score set
effect_ci_lower
Author Dataset EV3 75% ALDEx2 effect confidence-interval lower bound
effect_ci_higher
Author Dataset EV3 75% ALDEx2 effect confidence-interval upper bound
effect_fdr
Author Dataset EV3 ALDEx2 false-discovery rate
significant_fdr_0_05
True when effect_fdr is less than 0.05; this flag was not used to remove rows
effect_direction
Direction of the selected score; positive means lower protein abundance
rsid
dbSNP or gnomAD rs identifier when present in Dataset EV2
gnomad_allele_frequency
gnomAD v3.1.2 allele frequency when present
transcript
Transcript identifier from Dataset EV2
transcript_consequence
Transcript HGVS consequence from Dataset EV2
genomic_chromosome
GRCh38 chromosome from Dataset EV2
genomic_position
GRCh38 genomic position from Dataset EV2
clinvar_clinical_significance
ClinVar clinical-significance annotation from Dataset EV2
integrated_rna_effect
Dataset EV3 RNA-versus-gDNA ALDEx2 effect retained as cross-layer context
integrated_rna_fdr
Dataset EV3 RNA-versus-gDNA FDR
integrated_mid_polysome_effect
Dataset EV3 F3-versus-RNA ALDEx2 effect retained as cross-layer context
integrated_mid_polysome_fdr
Dataset EV3 F3-versus-RNA FDR
integrated_heavy_polysome_effect
Dataset EV3 F4-versus-RNA ALDEx2 effect retained as cross-layer context
integrated_heavy_polysome_fdr
Dataset EV3 F4-versus-RNA FDR
integrated_protein_effect
Dataset EV3 P3-versus-gDNA protein-abundance effect; positive means lower protein abundance
integrated_protein_fdr
Dataset EV3 protein-abundance effect FDR

Quality control

The authors called variants with satmut_utils v1.0.3-dev001, retained variants in the two mutagenesis ROIs, required amber UAG/TAG for nonsense calls, required log10 variant frequency greater than -5.8, and removed SNPs with false-positive random-forest predictions in more than half of the gDNA libraries. Variant frequencies were normalized to wild type with a 0.5 pseudocount, negative-control variants were subtracted, batches were ComBat-corrected, and technical replicates were summarized by the median. Only variants observed in all biological replicates and with median natural-log frequency greater than -10 were used for differential comparisons; low-yield/depth polysome fractions 1/2 and flow population 4 were excluded. ALDEx2 used four biological replicates. Package QC retained all 3050 published MAVEdb score rows with finite score and FDR values in the published score set; non-significant variants were retained, and unavailable source annotations are represented as NA.

Curation notes

The biological material is the HEK293T LLP iCasp9 Blast derivative; CVCL:0063 is the Cellosaurus identifier for the parent HEK293T line. MAVEdb HGVS coordinates are target-relative, so c.1 and p.1 correspond to the first assayed MB-COMT codon, Glu40; plasmid coordinates and variant_id preserve the author reference. The two ROIs are separated coding intervals on GRCh38 chr22. Cross-layer Dataset EV3 annotations and Dataset EV2 genomic annotations are included when available; missing values are NA. The MAVEdb score-set accession defines the retained records; selected effect values, intervals, and FDRs use the full-precision author Dataset EV3 layer and agree with MAVEdb up to API rounding. The other integrated effects are contextual annotations. Raw sequencing reads were not included.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.