Experiment / E1MFDQI7JEpisomal Plasmid MPRA

Primary K562 MPRA of platelet-function variants

Bayesian modelling of high-throughput sequencing assays with malacoda

An episomal plasmid MPRA tested 2,666 biallelic, platelet-function-related variant constructs using synthetic oligonucleotides with approximately 150 bp of genomic context and inert 14-bp barcodes. Three DNA input and six RNA output sequencing replicates were measured after transfection into K562 cells; the released primary data use anonymized variant identifiers.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

Synthetic oligonucleotides carrying approximately 150 bp of genomic context around the tested variants and inert 14-bp barcodes were cloned into a plasmid reporter library and transfected into K562 cells. The release contains MPRA_DNA1–3 and MPRA_RNA1–6 barcode counts, and primary_comparisons contains the authors’ transcription-shift estimates from malacoda marginal and conditional priors, MPRAscore, mpralm, a t-test, MPRAnalyze, and QuASAR-MPRA. The primary IDs and nucleotide sequences are anonymized or not released in S2 Data.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (37 of 37)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 37 definitions
anonymized_variant_id
Anonymous variant or construct identifier supplied in S2 primary_mpra_data and primary_comparisons.
raw_ref_barcodes
Number of released reference-allele barcode rows for the variant before DNA-representation QC.
raw_alt_barcodes
Number of released alternate-allele barcode rows for the variant before DNA-representation QC.
qc_ref_barcodes
Number of reference barcode rows remaining after the malacoda-style DNA-representation filter.
qc_alt_barcodes
Number of alternate barcode rows remaining after the malacoda-style DNA-representation filter.
low_qc_allele_representation
True when either allele has fewer than two retained barcode rows; these rows are retained but should be interpreted cautiously.
qc_status
Variant-level QC status; all emitted rows passed the biallelic and at-least-one-retained-barcode-per-allele filter.
ref_dna1
Sum of MPRA_DNA1 counts across QC-passing reference barcodes.
ref_dna2
Sum of MPRA_DNA2 counts across QC-passing reference barcodes.
ref_dna3
Sum of MPRA_DNA3 counts across QC-passing reference barcodes.
ref_rna1
Sum of MPRA_RNA1 counts across QC-passing reference barcodes.
ref_rna2
Sum of MPRA_RNA2 counts across QC-passing reference barcodes.
ref_rna3
Sum of MPRA_RNA3 counts across QC-passing reference barcodes.
ref_rna4
Sum of MPRA_RNA4 counts across QC-passing reference barcodes.
ref_rna5
Sum of MPRA_RNA5 counts across QC-passing reference barcodes.
ref_rna6
Sum of MPRA_RNA6 counts across QC-passing reference barcodes.
alt_dna1
Sum of MPRA_DNA1 counts across QC-passing alternate barcodes.
alt_dna2
Sum of MPRA_DNA2 counts across QC-passing alternate barcodes.
alt_dna3
Sum of MPRA_DNA3 counts across QC-passing alternate barcodes.
alt_rna1
Sum of MPRA_RNA1 counts across QC-passing alternate barcodes.
alt_rna2
Sum of MPRA_RNA2 counts across QC-passing alternate barcodes.
alt_rna3
Sum of MPRA_RNA3 counts across QC-passing alternate barcodes.
alt_rna4
Sum of MPRA_RNA4 counts across QC-passing alternate barcodes.
alt_rna5
Sum of MPRA_RNA5 counts across QC-passing alternate barcodes.
alt_rna6
Sum of MPRA_RNA6 counts across QC-passing alternate barcodes.
ref_dna_mean_cpm
Mean reference DNA count per million across the three DNA samples; sample depths use total released sample counts.
alt_dna_mean_cpm
Mean alternate DNA count per million across the three DNA samples; sample depths use total released sample counts.
ref_rna_mean_cpm
Mean reference RNA count per million across the six RNA samples; sample depths use total released sample counts.
alt_rna_mean_cpm
Mean alternate RNA count per million across the six RNA samples; sample depths use total released sample counts.
derived_log2_alt_ref_rna_dna_activity
Descriptive log2 alternate/reference activity shift: log2[((alt RNA CPM + 1)/(alt DNA CPM + 1))/((ref RNA CPM + 1)/(ref DNA CPM + 1))]. This is a derived summary, not a paper model output.
author_ts_malacoda_marginal
Authors’ marginal-prior malacoda transcription-shift estimate from primary_comparisons, on the paper’s natural-log activity scale.
author_ts_malacoda_conditional
Authors’ DeepSea-informed conditional-prior malacoda transcription-shift estimate from primary_comparisons, on the paper’s natural-log activity scale; blank when unavailable.
author_ts_mprascore
Authors’ MPRAscore transcription-shift estimate from primary_comparisons.
author_ts_mpralm
Authors’ mpralm transcription-shift estimate from primary_comparisons.
author_ts_t_test
Authors’ t-test transcription-shift estimate from primary_comparisons; blank when unavailable.
author_ts_mpranalyze
Authors’ MPRAnalyze transcription-shift estimate from primary_comparisons.
author_ts_quasar_mpra
Authors’ QuASAR-MPRA transcription-shift estimate from primary_comparisons.

Quality control

The paper’s malacoda workflow uses a default 15th-percentile DNA-representation cutoff, retaining barcodes whose mean depth-adjusted DNA count is above the cutoff. Recalculation from S2 primary_mpra_data gave a cutoff of 0 because of zero-inflation and retained 59,755 of 73,124 barcode rows (81.72%). All 2,666 biallelic variant IDs had at least one retained barcode for each allele and were retained in the variant-level table. The table aggregates only retained barcode rows; two variants have only one retained barcode for one allele and are flagged by low_qc_allele_representation. The raw count fields had no missing or negative values. Missing published conditional-prior and t-test estimates remain blank where the authors supplied missing values. The paper also excludes the largest 5% of dispersion estimates when fitting empirical priors; no new Bayesian fits were run for this package.

Curation notes

The single processed table is variant-level (2,666 rows) and was generated from the raw S2 RData file. Barcode-level sequences and counts remain available in raw_data/plos_supplement_S2_data.RData. The S2 release has 73,124 unique barcode rows (approximately 13–14 barcode rows per allele) and no nucleotide-level allele or genomic coordinate mapping, so rsIDs, reference/alternate bases, region coordinates, and a reference assembly cannot be assigned. The descriptive CPM/activity columns were calculated from QC-passing barcode sums, while the author_ts_* columns are copied from the authors’ primary_comparisons object and are not p-values or newly re-estimated effects. The S1 raw file retains the small rsID-labelled MPRA/luciferase validation subset, but it cannot be safely joined to the anonymous primary IDs without an author-provided mapping.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.